Xiaomi releases Xiaomi-Robotics-1 VLA model
Xiaomi has introduced Xiaomi-Robotics-1, a Vision-Language-Action robot foundation model pre-trained on over 100,000 hours of diverse real-world trajectories. The model learns general action-generation capabilities decoupling hardware from pre-training, achieving state-of-the-art simulation results and rapid physical robot adaptation.
Xiaomi-Robotics-1 shows that the scaling laws of vision and language models can successfully transfer to physical agents by decoupling action-generation learning from specific hardware configurations.
* Decoupling embodiment from initial pre-training solves the robotics data bottleneck by utilizing massive quantities of diverse video demonstration data.
* Strong scaling laws are confirmed: larger pre-trained models and larger datasets steadily lower action validation errors and translate directly to higher real-robot success rates.
* High data efficiency in downstream tasks allows robots to adapt to complex manipulations like laundry loading and phone packing with less than 10 hours of training data.
DISCOVERED
17h ago
2026-07-20
PUBLISHED
21h ago
2026-07-20
RELEVANCE
AUTHOR
ilreb