SCOPE-4D is an endoscopic 4D geometry foundation model that predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. The authors introduce SCOPE-5K, a curated dataset of about 5,000 real and synthetic gastrointestinal endoscopy and laparoscopy clips, and use it for geometric supervised fine‑tuning. Adding Common–Residual Motion (CRM) constraints and trajectory supervision further improves camera and depth estimation and enables dense 3D tissue tracking, as shown by evaluations on public and new benchmarks and a blinded user study.
By Chaoyi Zhou, Zhongpai Gao, Anwesa Choudhuri, Meng Zheng, Benjamin Planche, Run Wang, Terrence Chen, Siyu Huang, Ziyan Wu
arXiv:2610.03380v1 Announce Type: new
Abstract: Deepfake detection in videos remains challenging, as manipulated content may appear visually consistent at the frame level while exhibiting subtle temp...
By Chahira Benhama, Mohand Sa\"id Allili, Assia Hamadene
arXiv:2610.02697v1 Announce Type: cross
Abstract: Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instruct...
By Yixuan Jiang, Wentong Li, An Liu, Zihao Xin, Fulin Tang, Cong Leng, Yang Gao, Jian Cheng
arXiv:2610.03283v1 Announce Type: cross
Abstract: The ability of a device to localize itself within its surroundings is a fundamental prerequisite for spatial computing. Visual-inertial odometry (VIO...
By Patrick Wolf, Mateo de Mayo, Daniel Cremers
arXiv:2606.03915v2 Announce Type: replace
Abstract: We propose PatchScene, a novel diffusion-based framework for large-scale LiDAR scene completion. Unlike existing methods that rely on global latent...
By Qingdong Xu, Jiajun Zhu, Shilin Zhu, Xinjing He, Chao Lu, Huanran Wang, Jiyao Zhang
arXiv:2610.02338v1 Announce Type: cross
Abstract: A growing body of work suggests that tactile sensing gives robot policies contact information that complements vision in dexterous manipulation. Howe...
By Jingyun Yang, Baiyu Shi, Timothy Yu, Haitian Liu, Alberta Longhini, Weichen Wang, Rika Antonova, Zhenan Bao, Jeannette Bohg
arXiv:2610.02432v1 Announce Type: cross
Abstract: Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This t...
By Narek Maloyan
arXiv:2610.02339v1 Announce Type: cross
Abstract: Robot demonstrations may contain useful behavior even when individual episodes are inefficient or unsuccessful. Trajectory stitching offers a way to...
By Juntao Ren, Yifan Hou, Shuran Song
arXiv:2610.02334v1 Announce Type: cross
Abstract: Mobile embodied AI networks (MEAN) enable embodied agents to perceive, reason, communicate, and act in wireless environments. In such networks, agent...
By Yahao Ding, Jiaxiang Wang, Zhouxiang Zhao, Zhaohui Yang, Mingzhe Chen, Mohammad Shikh-Bahaei
arXiv:2610.02508v1 Announce Type: new
Abstract: World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an in...
By Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang, Pengfei Liu, Ya Zhang, Michal Drozdzal, Amir Bar
arXiv:2602.16543v2 Announce Type: replace
Abstract: Safe reinforcement learning (Safe RL) learns robotic controllers that optimize task rewards under safety constraints, yet observation perturbations...
By Jialiang Fan, Shixiong Jiang, Mengyu Liu, Fanxin Kong
arXiv:2610.02513v1 Announce Type: cross
Abstract: Large-scale vectorized HD maps provide structured road information that is essential for perception, localization, and planning in autonomous driving...
By Ziwei Li, Yi-Tang Chen, Xiaoqi Wang, Wenbin He, Han-Wei Shen, Liu Ren
arXiv:2610.03717v1 Announce Type: cross
Abstract: This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene struc...
By Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg
arXiv:2610.03476v1 Announce Type: cross
Abstract: Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and...
By Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi
arXiv:2209.14887v3 Announce Type: replace-cross
Abstract: Robotic locomotion is often approached with the goal of maximizing robustness and reactivity by increasing motion control frequency. We chall...
By Siddhant Gangapurwala, Luigi Campanaro, Ioannis Havoutis
arXiv:2608.25572v2 Announce Type: replace-cross
Abstract: Action-conditioned world models have become an important foundation for embodied prediction, planning, and synthetic data generation, but the...
By Xiang Liu, Kunwei Wu, Miao Liu, Sen Cui, Changshui Zhang
arXiv:2608.05369v2 Announce Type: replace-cross
Abstract: Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct r...
By Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu
arXiv:2609.20822v2 Announce Type: replace-cross
Abstract: Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and age...
By Bingxin Xu, Yuzhang Shang, Zhen Dong, Emilio Ferrara
arXiv:2610.03695v1 Announce Type: new
Abstract: Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language mod...
By Adithya Bhaskar, Jeffrey Cheng, Danqi Chen
arXiv:2610.03710v1 Announce Type: cross
Abstract: Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. E...
By Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik, C. Karen Liu, Ken Goldberg, Angjoo Kanazawa