Wiggle and Go! is a two‑stage framework for zero‑shot rope manipulation that first performs a brief, safe wiggle action to infer rope parameters, then uses those parameters to condition a trajectory optimizer for goal‑conditioned execution. The method achieves 3.55 cm average accuracy on 3D target striking in real‑world tests, far outperforming uninformed baselines, and secures over 50% success on multi‑objective lobbing and draping tasks. Predicted parameters transfer well to unseen motions, with a 0.95 Pearson correlation between simulated and real rope dynamics, demonstrating task‑agnostic generalization without retraining.
By Arthur Jakobsson, Abhinav Mahajan, Karthik Pullalarevu, Krishna Suresh, Yunchao Yao, Yuemin Mao, Bardienus Duisterhof, Shahram Najam Syed, Jeffrey Ichnowski
arXiv:2609.25836v1 Announce Type: new
Abstract: Multi-task optimization (MTO) addresses a set of optimization tasks simultaneously, often suffering from inaccurate inter-task relationship estimation...
By Tingyang Wei, Haofeng Wu, Jiao Liu, Zhao Wei, Puay Siew Tan, Yew-Soon Ong
The paper investigates how Vision‑Language‑Action (VLA) models can generalise across different driving environments and camera setups. It introduces a multi‑dataset training strategy and an auxiliary objective called BEV‑Forcing, which injects bird‑eye‑view spatial information into the VLA backbone to improve both in‑distribution and out‑of‑distribution performance on a limited number of camera rigs. The authors observe that while BEV‑Forcing helps when training data is scarce, its advantage diminishes as the number of training embodiments grows, suggesting that scaling diversity may reduce the impact of such auxiliary tasks.
By Caio Azevedo, Stefano Sabatini, Sascha Hornauer, Fabien Moutarde
arXiv:2601. 19810v2 Announce Type: replace-cross Abstract: Unsupervised pre-training can equip reinforcement learning agents with prior knowledge and accelerate learning in downstream tasks.
By Octavio Pappalardo
arXiv:2603. 25464v2 Announce Type: replace-cross Abstract: Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies for any reward function directly at test time.
By Jiajun Hu, Nuria Armengol Urpi, Jin Cheng, Stelian Coros
arXiv:2602. 01962v2 Announce Type: replace-cross Abstract: Off-policy learning methods seek to derive an optimal policy directly from a fixed dataset of prior interactions.
By Arip Asadulaev, Maksim Bobrin, Salem Lahlou, Dmitry Dylov, Fakhri Karray, Martin Takac