The paper introduces a fast Bayesian method for estimating homographies from noisy point correspondences, providing a posterior distribution over the homography parameters. A closed‑form solution for the posterior mean in homogeneous coordinates is derived, complemented by an iterative Bayesian approach to address non‑linearities. Experiments on synthetic data and real image stitching show improved accuracy over DLT and supply uncertainty estimates for the homography.
By Hanne Beuter, Sebastian Dorn
arXiv:2609.14261v1 Announce Type: cross
Abstract: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative...
By Prajwal Koirala, Mark Campbell
arXiv:2609.13352v1 Announce Type: cross
Abstract: We present three robotic art installations which explore the aesthetics of adaptive behavior. Through embodied machine leaning and digital evolution,...
By Sofian Audry, Stephen Kelly
arXiv:2609.14313v1 Announce Type: cross
Abstract: Robust point tracking in endoscopic videos is essential for computer-assisted intervention and autonomous robotic surgery, enabling continuous regist...
By Jiaming Zhang, Zijian Wu, Mehran Armand, Septimiu Salcudean
arXiv:2609.15361v1 Announce Type: cross
Abstract: Effective communication is a cornerstone of distributed intelligence in Multi-Agent Reinforcement Learning (MARL), yet ensuring that generated messag...
By Rafael Pina, Varuna De Silva, Corentin Artaud
The paper introduces a ROS-Agent architecture that enhances task reliability and execution efficiency for open‑source LLM‑powered robotic agents. It adds a MetaTool that forces the LLM to produce a structured pseudo‑code plan before any action, storing this plan in a scratchpad to separate planning from execution. Experiments on a custom mobile robot show up to ~24% improvement in complex task completion and contextual consistency compared to the baseline.
By Kazi Abrar Mahmud, Nilotpaul Kundu Dhurubo, Tamal Kirttonia, Sabbir Hossain Ujjal, Mohammad Ariful Haque
arXiv:2609.15113v1 Announce Type: cross
Abstract: As robotic systems grow more general, legal norms are needed to integrate them into society. This paper extends the isomorphism problem of aligning l...
By Dylan Waldner, Yiannis Kantaros, Guido Governatori, Risto Miikkulainen, Amir Banifatemi
arXiv:2609.13545v1 Announce Type: cross
Abstract: Learning-based adaptive control of robotic manipulators with non-observable friction memory has been addressed by attention- based meta-controllers w...
By Giansalvo Cirrincione, Adriano Fagiolini
arXiv:2609.13243v1 Announce Type: cross
Abstract: We present GzDRL, a novel single-process reinforcement learning (RL) framework for Gazebo that overcomes longstanding bottlenecks in scalable, reprod...
By Amal Dev Haridevan, Junjie Kang, Jinjun Shan
arXiv:2609.13295v1 Announce Type: cross
Abstract: Diffusion learning leverages the statistical mechanism of diffusion processes for learning, reasoning, and inferring complex distributions from data....
By Max Muchen Sun, Cem Bilaloglu, Ananya Rao, Stefan Ivic, Guillaume Sartoretti, Kathleen Fitzsimons, Ian Abraham, Sylvain Calinon, Todd Murphey
arXiv:2609.13152v1 Announce Type: new
Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experim...
By Joseph Chan, Utkarsh Jha, Xiyin Yang, Abhinav Jarajapu, Anik Sahai, Eddie Hu, Robin Jeshua Deepak, Stefano Saravalle, Aditya Shah
arXiv:2602.08370v2 Announce Type: replace-cross
Abstract: Realizing versatile and human-like performance in high-demand sports like badminton remains a formidable challenge for humanoid robotics. Unl...
By Yeke Chen, Shihao Dong, Xiaoyu Ji, Jingkai Sun, Zeren Luo, Liu Zhao, Jiahui Zhang, Wanyue Li, Ji Ma, Bowen Xu, Yimin Han, Xuanyi Li, Yudong Zhao, Liyun Li, Peng Lu
arXiv:2609.15014v1 Announce Type: cross
Abstract: Pretrained generative robot policies can produce effective behaviors across diverse environments, but deployment can lead to requirements and prefere...
By Yixuan Jia, Jonathan P. How
arXiv:2603.12916v4 Announce Type: replace-cross
Abstract: Multivariate time series anomalies often manifest as shifts in cross-channel dependencies rather than simple amplitude excursions. In autonom...
By Kadir-Kaan \"Ozer, Ren\'e Ebeling, Markus Enzweiler
The paper presents a method for generating lower‑limb joint‑angle gait trajectories using conditional diffusion models. It compares a baseline transformer diffusion model with a controllable diffusion transformer that includes adaptive normalization and classifier‑free guidance. Experiments on 4,590 gait cycles demonstrate that these diffusion models can produce realistic, periodic gait patterns while allowing some control over gait characteristics such as step length.
By Damian Benasco, Juan Carballeira-Lopez, Jaime Ramos-Rojas, Julio S. Lora-Millan, Antonio J. Del-Ama, David Rodriguez-Cianca, Pablo Lanillos
arXiv:2609.07534v2 Announce Type: replace-cross
Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sus...
By Yuhan Wang, Yurou Chen, Hongye Jiang, Wenzhao Lian
arXiv:2603.19308v2 Announce Type: replace-cross
Abstract: In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key...
By Wentao Wang, Haoran Xu, Guang Tan
arXiv:2609.13265v1 Announce Type: cross
Abstract: Near-infrared (NIR) imaging does not consistently outperform standard color cameras for daytime agricultural traversability once spatial data leakage...
By Sungwoo Kang
arXiv:2609.12606v2 Announce Type: replace
Abstract: While multimodal reasoning has advanced rapidly, solving complex geometry problems critically hinges on active visual assistance, such as construct...
By Zhitong Dong, Jicai Pan, Yingguo Gao, Jingting Ding, Hao Chen, Jinjie Gu
PhysBrain 1.5 is a unified vision‑language model that learns to understand physical environments, generate actions, and predict future states by encoding language, end‑effector motion, and dense visual targets as discrete sequences and training them with autoregressive next‑token prediction. The model is pre‑trained on human interaction videos and fine‑tuned on human demonstrations, robot trajectories, and simulated experience, achieving an average score of 72.5 across 28 embodied understanding benchmarks and outperforming other open‑source models on 14 of them. It also demonstrates the ability to produce end‑effector trajectories and predict future scenes with spatially aligned RGB, depth, and robot‑mask outputs.
By DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang