In-Context Robot Learning with VLM Agents
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.10522v1 Announce Type: cross Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains cha...
Teach-and-Grow Learning (TGL) is an agent-centered architecture that transforms a few successful demonstrations into reusable Skill Blocks, enabling a robot to compose, execute, and revise behaviors in new scenes without task-specific policy retraining. The system maintains a Skill Library and structured Experience Memory to capture successes, failures, and repairs, allowing persistent reuse and agent-directed adaptation. Evaluation on the LIBERO benchmark shows state-of-the-art performance, and the authors propose a scaling-law hypothesis suggesting that accumulated reusable experience reduces future-task error and teaching demand following a power-law trend.
arXiv:2601. 20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task-specific demonstrations and fine-tuning, and often generalize poorly under domain shift.
arXiv:2609.15683v1 Announce Type: new Abstract: While In-Context Learning (ICL) enables models to adapt from exemplars without parameter updates, multimodal ICL remains largely underexplored, particu...
Zero-WAM introduces a causal video-action model that enables robots to perform unseen manipulation tasks by following in-context human video guidance. The authors create HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks, and propose an in-context future chunk prediction objective to prevent shortcut learning. In simulation, Zero-WAM attains a 47.0% success rate on seven unseen tasks, outperforming the best video-action baseline by 29.5 percentage points, and demonstrates real‑world generalization to complex, long‑horizon, and fine‑grained tasks.
arXiv:2607. 23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail.