
Researchers have introduced GPT-Policy, a framework that enables robots to learn new manipulation skills in-context from human demonstrations and feedback without updating model weights.
GPT-Policy is an embodied AI framework designed to achieve in-context robot learning using general-purpose vision-language models (VLMs) without fine-tuning or parameter updates. By processing varied context sources—such as human video demonstrations, recorded robot teleoperations, goal images, and interactive feedback—the system enables a robot to interpret a new task and generate executable, verifiable trajectories from arbitrary initial states. This approach sidesteps the costly, data-hungry retraining cycles traditionally required for robot policy adaptation, paving the way for adaptable generalist robots that learn on the fly during deployment.
In-context learning for physical manipulation represents a major paradigm shift that could bypass brittle per-task fine-tuning loops in robotics.
- –Zero-shot skill acquisition: Eliminating gradient updates allows robots to generalize across novel manipulation tasks immediately after observing a single demonstration or prompt.
- –Multimodal context integration: Seamlessly unifies human videos, robot history, and corrective feedback into a coherent prompt representation for action generation.
- –Latency and compute constraints: Relying on heavy, frontier VLMs poses real-world latency challenges for high-frequency dynamic control loops.
- –Grounding and execution reliability: While high-level reasoning excels in-context, translating visual understanding into millimetric physical precision still requires robust low-level verification.
DISCOVERED
1h ago
2026-09-20
PUBLISHED
1h ago
2026-09-20
RELEVANCE
AUTHOR
AI Search