arXiv:2504. 17939v2 Announce Type: replace-cross Abstract: We present a computational model of the mechanisms that may determine infant behavior in the "mobile paradigm".
By Josua Spisak, Sergiu Tcaci Popescu, Stefan Wermter, Matej Hoffmann, J. Kevin O'Regan
arXiv:2608. 14144v1 Announce Type: cross Abstract: Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest.
By Yijiang Li, Yijun Liang, Yunjie Tian, Bingyang Wang, Ke Zhang, Zhenfei Yin, Di Fu, Philip Torr, Nuno Vasconcelos
arXiv:2609.23352v1 Announce Type: new
Abstract: Bimanual interaction produces complementary tactile views of the same physical process, yet existing tactile representation learning largely models the...
By Chenxin Liang, Youchen Lai, Chuqiao Lyu, Tianxing Chen, Shoujie Li, Wenbo Ding
OPD‑Aha is a privileged on‑policy distillation method that improves multimodal reasoning by reconstructing the distillation target from the teacher’s isolated visual preference instead of relying on fragile teacher‑student discrepancies. It suppresses continuations that contradict the image, encouraging students to interrupt flawed reasoning with reflection tokens such as "wait" and "actually." This approach leads to consistent improvements across fine‑grained perception and complex multimodal reasoning benchmarks.
By Chenhao Qiu, Dawei Li, Yechao Zhang, Lei Gong, Zhen Tan
arXiv:2609.24976v1 Announce Type: cross
Abstract: Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple pre...
By Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan
arXiv:2607. 23899v1 Announce Type: cross Abstract: This exploratory study examines whether a large multimodal language model, GPT-5.
By Roberto Spinelli, Thiago C. Martins