arXiv:2609.38443v1 Announce Type: cross
Abstract: We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yi...
By Cameron Smith, Arsh Tangri, Vitor Guizilini, Yue Wang, Zubair Irshad, Sergey Zakharov
arXiv:2609.38494v1 Announce Type: cross
Abstract: Robotic manipulation integrates vision, touch, and language, whose importance shifts across stages: vision guides reaching, while touch, through its...
By Amir-Hossein Shahidzadeh, Seungjae Lee, Eadom Dessalene, Shanthosh Raaj Mohanram Mageswari, Soroush Etemad, Furong Huang, Cornelia Ferm\"uller, Yiannis Aloimonos
arXiv:2609.38616v1 Announce Type: cross
Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...
By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin
arXiv:2609.39375v1 Announce Type: cross
Abstract: A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests...
By Hyunjoon Lee, Haebeom Jung, Eunsung Cha, Daeun Lee, Yu-Chiang Frank Wang, Jaesung Choe, Jaesik Park
arXiv:2609.39388v1 Announce Type: cross
Abstract: Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in acti...
By Wei Xue, Keliang Liu, Mingzhang Cui, Jinhua Xie, Jinjie Wei, Jianan Hou, Jingcheng Lu, Lintao Wang, Kaixiang Qiu, Yizhou Liu, Xinghai Ye, Jinghang Han, Mingcheng Li, Jie Gu, Shunli Wang, Lihua Zhang, Dingkang Yang
arXiv:2609.39324v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. Howeve...
By Jingqiu Wang, Yan Wang
arXiv:2601.08355v3 Announce Type: replace
Abstract: Vision-Language Models (VLMs) are increasingly deployed in autonomous driving and embodied AI systems, where reliable perception is critical for sa...
By Guo Cheng, Huang Li
arXiv:2604.18484v2 Announce Type: replace
Abstract: Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from co...
By Kangan Qian, ChuChu Xie, Yang Zhong, Jingrui Pang, Siwen Jiao, Sicong Jiang, Zilin Huang, Yunlong Wang, Kun Jiang, Mengmeng Yang, Hao Ye, Guanghao Zhang, Hangjun Ye, Guang Chen, Long Chen, Diange Yang
arXiv:2607.03470v2 Announce Type: replace
Abstract: Synthesizing physically accurate mirror reflections remains a fundamental challenge for modern text-to-image diffusion models, which are increasing...
By Xuan-Bach Mai, Duy-Phuc Nguyen, Quoc-Van Le, Tam V. Nguyen, Thanh-Toan Do, Huu Le, Duong-Van Nguyen, Minh-Triet Tran, Trung-Nghia Le
arXiv:2609.34220v2 Announce Type: replace-cross
Abstract: Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as ob...
By Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu, Pengfei Liu, Jiarui Zhang, Fangqiang Ding, Lihua Xie, Gen Li, Jianfei Yang
arXiv:2609.38855v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic m...
By Gongxin Yao, Yongsheng Zhao, Jiayin Deng, Deng Liang, Han Gao, Lei Zhao, Baoping Cheng
arXiv:2607.27627v2 Announce Type: replace-cross
Abstract: Unmanned aerial vehicle (UAV) relay networks can restore connectivity after communication infrastructure is damaged. Urban relay placement is...
By Dohun Lee, Kyeonghyun Yoo, Seokmin Kim, Byongho Lee, Seungjoo Oh, Hwangnam Kim
arXiv:2609.38984v1 Announce Type: cross
Abstract: World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and...
By Xinling Xie, Haodong Wang, Jiazhi Mi, Zhiming Liu, Zicong Hong, Xiaoyi Pang, Qianli Liu, Yangjia Hu, Ying Chen, Zhengyang Yan, Song Guo
arXiv:2609.40031v1 Announce Type: cross
Abstract: Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated co...
By Khaled Abud, Aleksey Yakushev, Aleksandr Akimenkov, Irina Serzhenko, Kirill Aistov, Egor Kovalev, Dmitry Obydenkov, Sergey Lavrushkin, Anastasia Antsiferova, Dmitriy Vatolin, Yury Markin, Kirill Lukianov
arXiv:2609.39665v1 Announce Type: new
Abstract: Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This...
By Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong
arXiv:2609.38251v1 Announce Type: cross
Abstract: The rapid evolution of image manipulation techniques has raised growing public security concerns. Existing Image Forgery Localization (IFL) methods c...
By Chenqi Kong, Song Xia, Anwei Luo, Peisong He, Alex C. Kot, Yuming Fang
arXiv:2609.38615v1 Announce Type: cross
Abstract: Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Ex...
By Hongjia Zhai, Xiyu Zhang, Haoran Zhang, Zhichao Ye, Haomin Liu, Guofeng Zhang, Ian Reid, Xingxing Zuo
arXiv:2609.38733v1 Announce Type: new
Abstract: Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decis...
By Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, Samuel J. Gershman
arXiv:2609.38463v1 Announce Type: cross
Abstract: Autonomous driving planners are typically evaluated using aggregate metrics such as driving score, destination rate, and collision rate, which do not...
By Victoria Smirnova, Viktoriia Zinkovich, Gregorii Bukhtuev, Artem Belyaev, Andrey Kuznetsov, Denis Shepelev, Vlad Shakhuro
arXiv:2609.39178v1 Announce Type: cross
Abstract: Recently, Vision-Language-Action (VLA) models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understa...
By Songhua Yang, Ziyu Liu, Yuanwei Liu, Xuetao Li, Xuanye Fei, He Huang, Zheng Wang, Miao Li