arXiv:2609.33965v2 Announce Type: replace-cross
Abstract: We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between i...
By Joshua Horswill, Ross Hunter, Matt Clifford, James Hall
arXiv:2609.36691v1 Announce Type: new
Abstract: Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embod...
By Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu, Haoran Lu, Yijiang Li, Yaqi Xie, Katia P. Sycara, Han Liu
arXiv:2609.35916v1 Announce Type: cross
Abstract: Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions fac...
By Jie Yang, Jiajun Chen, Jiazheng Zhou, Mianqiu Huang, Yining Zheng, Yuxin Wang, Xipeng Qiu
arXiv:2609.37476v1 Announce Type: cross
Abstract: Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environme...
By Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu, Zhengcheng Shen, Harold Soh
arXiv:2609.35823v1 Announce Type: new
Abstract: Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes....
By Kang Yang, Shuai Liu, Hang Li, Yance Fang, Deying Li, Yongcai Wang
arXiv:2609.36454v1 Announce Type: new
Abstract: We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanic...
By Wenliang Guo, Zhanbo Huang, Yu Kong
arXiv:2609.36755v1 Announce Type: new
Abstract: Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of...
By Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou
arXiv:2609.36801v1 Announce Type: new
Abstract: Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, i...
By Minkwan Kim, Junho Kim, Seungmin Lee, Changwoon Choi, Young Min Kim
arXiv:2609.36844v1 Announce Type: new
Abstract: Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the b...
By Suhani Grover, Astik Srivastava, Viswas Dinesh, Avinash Sharma, K. Madhava Krishna
arXiv:2609.37089v1 Announce Type: new
Abstract: Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, exec...
By Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu, Bo Dai, Mulin Yu, Tao Lu
arXiv:2609.37135v1 Announce Type: new
Abstract: Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-g...
By Yi-Pei Chiu, Wei-Ta Chu
arXiv:2609.37495v1 Announce Type: new
Abstract: Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While exis...
By Yun Chen, Munchurl Kim, Jeonghyeok Do
arXiv:2609.37655v1 Announce Type: new
Abstract: Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA)...
By Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu, Daoguo Dong, Jiachen Xu, Xin Tan
arXiv:2609.37870v1 Announce Type: new
Abstract: Raindrops adhered to camera lens or windshield are inevitable in rainy scenes and can become an issue for many computer vision systems such as autonomo...
By Zhixiang Hao, Shaodi You, Yu Li, Kunming Li, Feng Lu
arXiv:2609.38057v1 Announce Type: new
Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...
By Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
arXiv:2609.38163v1 Announce Type: new
Abstract: World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and pred...
By Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang
arXiv:2609.36520v1 Announce Type: cross
Abstract: RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety laye...
By Seungyeon Yoo, Gawon Lee, Seungwoo Jung, Inkyu Jang, H. Jin Kim
arXiv:2609.36779v1 Announce Type: cross
Abstract: Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker ac...
By Xiaotian Zhang, Yusheng Wang, Naoya Kagawa, Noritaka Takamura, Keiji Okuhara, Hiroyasu Baba, Jun Ota
arXiv:2609.37602v1 Announce Type: cross
Abstract: Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where envi...
By Michele Antonazzi, Alejandra C. Hernandez, Jos\'e Araujo, Olov Andersson, Patric Jensfelt
arXiv:2609.38059v1 Announce Type: cross
Abstract: Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a sca...
By Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia, Haoyang Huang, Nan Duan, Jiaya Jia