arXiv:2608. 11204v1 Announce Type: cross Abstract: Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.
By Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang
arXiv:2605.08712v2 Announce Type: replace
Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dime...
By Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang, Youqi Tao, Junfeng Duan, Daguang Xu, Qi Dou, Xin Jin, Wenjun Zeng, Hao Zhao, Yueming Jin
arXiv:2609.19142v1 Announce Type: new
Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse...
By Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski
arXiv:2608. 20284v1 Announce Type: new Abstract: Reliable surgical planning requires models to anticipate not only how instruments will move, but also how the operative visual state will evolve together with such motion.
By Weiliang Huang, Huanrong Liu, Bob Zhang, Qi Dou, Zhen Chen, Yun Gu, Guy Rosman, Qingbiao Li
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into do...
arXiv:2606. 10025v1 Announce Type: cross Abstract: We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution.
By Sriram Krishna, Ben Eisner, Haotian Zhan, Ying Yuan, Haoyu Zhen, Chuang Gan, Shubham Tulsiani, David Held
This survey reviews recent advances in surgical video generation, categorizing methods into unconditional, conditional, and world modeling generation. It highlights a shift from creating visually plausible frames to modeling the causal dynamics of surgical scenes, and discusses challenges such as pixel-level fidelity versus clinical plausibility, generalization, physical realism, controllability, and interpretability. The paper also compiles experimental results from public datasets to serve as a quantitative benchmark for the field.
By Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang
arXiv:2609.38059v1 Announce Type: cross
Abstract: Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a sca...
By Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia, Haoyang Huang, Nan Duan, Jiaya Jia
Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations.
arXiv:2607. 10706v1 Announce Type: cross Abstract: The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions.
By Haojie Huang, Zhang Ye, Linfeng Zhao, Boce Hu, Mingxi Jia, Yu Qi, Ahmed Agha, Dian Wang, Robert Platt, Robin Walters
arXiv:2606. 13769v1 Announce Type: cross Abstract: World models that capture how actions induce physical change enable scalable robot learning without reliance on embodiment-specific action labels.
By Seungjae Lee, Yoonkyo Jung, Jusuk Lee, Jonghun Shin, Amir Hossein Shahidzadeh, Yao-Chih Lee, H. Jin Kim, Jia-Bin Huang, Furong Huang
arXiv:2505. 03296v2 Announce Type: replace-cross Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy representation and imitation learning in robot manipulation.
By Jan Ole von Hartz, Adrian R\"ofer, Joschka Boedecker, Abhinav Valada