arXiv:2609.37907v1 Announce Type: new
Abstract: Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrati...
By Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio
arXiv:2609.36250v1 Announce Type: new
Abstract: Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many exi...
By Sanghyun Hahn, Jonghyun Choi
arXiv:2609.36172v1 Announce Type: cross
Abstract: Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot i...
By Saptak Das, Monidipa Das
arXiv:2609.36352v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multip...
By Ziyi Yin, Sangmin Woo, Kang Zhou, Sungyeon Kim, Aosong Feng, Haibo Ding, Jun Huan
arXiv:2609.37070v1 Announce Type: cross
Abstract: Rare but consequential failures can persist in learned locomotion policies for legged robots even when average task performance is high, in part beca...
By Ivan Ovinnikov, Pascal Sutter, Christian Gehring, Jordis Herrmann
arXiv:2609.37330v1 Announce Type: cross
Abstract: Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators...
By Thomas Deixelberger, Markus Steinberger
arXiv:2609.37359v1 Announce Type: cross
Abstract: Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves ou...
By Yifan Kang, Zihan Wang, Zhiwen Fan, Bangya Liu
The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.
By Ying Yuan, Yu Wang, Yize Cheng, Xuyang Wu
arXiv:2609. 03842v2 Announce Type: replace Abstract: Behavior regularization in offline reinforcement learning limits the exploitation of critic errors, but strong anchoring can also restrict policy improvement.
By Soohyun Choi, Seonvin Cho, Songnam Hong
The paper introduces Alignment‑Guided Flow Transformer (AGFT), a framework for Vision‑Language‑Action (VLA) models that explicitly enforces tri‑modal alignment among vision, language, and action through a dedicated alignment loss. AGFT bridges representational gaps across modalities, improving task adaptation and robustness, and employs a flow‑matching objective to reduce inference steps compared to diffusion‑based policies. Experiments on a large benchmark demonstrate that AGFT achieves higher success rates and lower inference latency than state‑of‑the‑art baselines, highlighting tri‑modal alignment as crucial for scalable VLA manipulation.
By Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao
arXiv:2603.12717v2 Announce Type: replace-cross
Abstract: Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are de...
By Tuan Duong Trinh, Basim Azam, Mohammed Ishaq Ansari, Mohammed Yaqoob Ansari, Naveed Akhtar
arXiv:2606.16996v2 Announce Type: replace-cross
Abstract: Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabu...
By Tran Dinh Tien, Zhiqiang Shen
OVIG is an optimistic verification framework that audits AI training by replaying the process and comparing gradient differences against an empirically calibrated boundary. It treats any gradient difference exceeding this boundary as a malicious deviation. By partitioning training into stride‑s intervals and storing evidence only at interval endpoints, OVIG dramatically reduces off‑chain storage and transmission costs while maintaining zero attack success rate across language, vision, and diffusion workloads.
By Hongxu Su, Jianzhu Yao, Huan Zhang, Xuechao Wang, Pramod Viswanath
arXiv:2609.34911v2 Announce Type: replace-cross
Abstract: Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replann...
By Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye
arXiv:2609.37098v1 Announce Type: cross
Abstract: Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providin...
By Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua, Haotian Shi, Wei Zhang, Lin Wang, Bin Ran
Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geome...
Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a bett...
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effe...
Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated...
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of lo...