Xiaomi Demonstrates Scaling Laws Apply to Robotics Policy Models
Key Takeaways
- ▸Xiaomi-Robotics-1 proves scaling laws apply to robot policy models, with real-world task success improving predictably as data and model scale increase
- ▸Embodiment-free (UMI) pre-training on 100,000 hours of diverse manipulation data breaks through robotics' data scarcity bottleneck
- ▸Automated vision-language-model-powered annotation enables large-scale labeling of manipulation trajectories without manual effort
Summary
Xiaomi has introduced Xiaomi-Robotics-1, a foundation model that demonstrates scaling laws—long proven effective for language and vision—now drive progress in robotics. The model addresses robotics' critical data scarcity bottleneck through a two-stage training approach: large-scale embodiment-free (UMI) pre-training on 100,000 hours of diverse manipulation trajectories from household, commercial, and industrial environments, followed by post-training on 7,200+ hours of real-robot data collected in actual homes. Using an automated annotation pipeline powered by vision-language models, Xiaomi labeled real-world manipulation videos with state-transition descriptions, enabling the model to learn generalizable action generation at unprecedented scale.
The research demonstrates clean scaling behavior across both pre-training and real-world deployment. Validation action error decreases predictably as model size and pre-training data increase, and critically, these scaling benefits transfer to actual robot performance in unseen environments with novel objects. The model can be deployed out-of-the-box to perform a wide range of mobile manipulation tasks, with success rates rising steadily as the underlying pre-trained model grows stronger. This work suggests that foundation model scaling laws—the principles that have driven LLM and vision model capabilities—can finally unlock comparable advances in robotics.
- Performance gains from pre-training scale transfer successfully to real robots in unseen environments, validating the approach's generalization ability
Editorial Opinion
Xiaomi's demonstration that scaling laws work for robotics represents a potentially transformative moment for the field. Robotics has long trailed language and vision models precisely because acquiring large-scale, high-quality training data from physical robots is prohibitively expensive. The clever use of embodiment-free (UMI) data sidesteps this constraint and unlocks foundation-model-scale training. The real test lies ahead: whether this approach can scale economically to production systems and generalize across the full spectrum of real-world tasks. If Xiaomi clears that hurdle, we may finally see robotics experience the capability step-change that language and vision models have already delivered.


