Show-Harness lets foundation VLMs control physical robots
Show-Harness is an open-source framework developed by Show Lab at the National University of Singapore that connects vision-language models directly to robotic control systems. By translating discrete semantic action units into deterministic metric motions across diverse robot platforms, it enables zero-shot manipulation with frontier VLMs and high-frequency execution on lightweight models with minimal fine-tuning.
Specialized, multi-billion-parameter vision-language-action models may be overcomplicating embodied intelligence by forcing neural networks to regress low-level motor metrics end-to-end; establishing a clean, symbolic interface demonstrates that foundation VLMs already possess sufficient spatial and physical reasoning to operate robots out of the box. Offloading metric translation to deterministic hardware interpreters decouples abstraction layers, allowing the VLM to focus purely on high-level spatial perception and sequential decision-making. Fine-tuning lightweight open-source backbones requires only a few GPU-hours while achieving control speeds between 12 and 33 Hz, matching or exceeding dedicated models like π₀.₅ and GR00T. Furthermore, the parameter-free semantic action space enables cross-embodiment generalization across various robot arms without architectural changes, while the GUMI interface replaces cumbersome teleoperation rigs with standard graphical interfaces to simplify data collection.
DISCOVERED
57m ago
2026-09-13
PUBLISHED
1h ago
2026-09-13
RELEVANCE
AUTHOR
AI Search