Skild AI S1 turns videos into robot skills
Skild AI’s S1 uses a single video demonstration as an in-context prompt for unseen, long-horizon manipulation tasks without task-specific fine-tuning. The company reports runs lasting up to 10 minutes and 66% per-step success on unseen tasks versus 9% for language prompting.
S1’s breakthrough is shifting robot adaptation from repeated data collection toward demonstration-and-run deployment. The results are promising, but remain company-reported internal evaluations rather than proof of reliable end-to-end autonomy.
- –Video preserves action order, geometry, and timing that language instructions often underspecify.
- –At 100,000 training hours, S1 reached 66% average per-step success on unseen tasks versus 9% for a language-prompted baseline.
- –Skild estimates one video prompt matched roughly 380 post-training demonstrations, though conventional post-training eventually reached 86% with 2,000 examples.
- –The benchmark used human intervention to recover from failures, so developers should treat the figures as research signals rather than production reliability guarantees.
DISCOVERED
1h ago
2026-08-30
PUBLISHED
1h ago
2026-08-30
RELEVANCE
AUTHOR
AI Search