Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal...
Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, das...
Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning ho...
Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in...
The paper introduces Trust Guided Decision Transformer (TGDT), a method that mitigates performance degradation in Decision Transformers during long rollouts by monitoring the model’s next‑state prediction error. TGDT evaluates multiple recent context suffixes, filters out those whose prediction error exceeds a calibrated threshold, and then selects the highest‑value action from the remaining trusted suffixes using a frozen critic. Experiments on D4RL navigation and locomotion tasks demonstrate that TGDT reduces persistent high‑error runs and improves returns compared to vanilla Decision Transformer and other context‑control baselines.
By Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath
The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.
By Mehmet Kerem Turkcan
The paper introduces Morphogene, a compact latent blueprint that links an agent’s body and control policy, enabling coordinated changes at the limb level through AdaConcat. Building on this, GeCode treats co-design as exploration within Morphogene space, using local refinement and global exploration to efficiently navigate design regions. Experiments on various 2D and 3D tasks show that GeCode outperforms state‑of‑the‑art methods, achieving faster convergence and higher performance.
By Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng
The paper introduces the Neural Action Codec (NAC), a convolutional encoder‑decoder architecture that treats short robot action trajectories as multi‑channel 1D signals and compresses them using a multi‑scale residual vector quantization (RVQGAN) model. NAC replaces traditional discrete action tokenizers with a compact, ordered token space via offset codebooks, allowing standard autoregressive policies to operate over short, structured sequences while a Vocos‑style decoder reconstructs the actions. Experiments on LIBERO‑10, RoboMimic, and real‑world manipulation tasks show that NAC achieves higher reconstruction fidelity and better average success rates than existing binning, FAST, and VQ‑based tokenizers at comparable or improved compression rates.
By Ahad Jawaid, Yu Xiang
The paper introduces an online model‑based reinforcement learning framework that learns a probabilistic dynamics ensemble from scratch for sampling‑based model predictive control, specifically targeting precise, high‑speed control of hydraulic excavators. A precision‑gated contouring objective prioritizes path accuracy over speed, enabling the system to achieve higher sample efficiency than existing model‑based RL baselines in a data‑driven simulator. The method is validated on an 11.5‑ton Menzi Muck M445 excavator, reaching tracking accuracy comparable to prior controllers after only 20 minutes of real‑world interaction and maintaining sub‑centimeter mean path error at high speeds after 40 minutes.
By Claudio Canales, Fang Nan, Marco Hutter, Javier Ruiz-del-Solar
Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support t...
The paper introduces P2P‑T, a data‑efficient, object‑centric framework that learns tool manipulation directly from human video demonstrations. It uses a two‑stage approach: first pretraining an object‑centric world model to extract stable pose priors, then integrating these priors into a pose‑aware low‑level policy. By automating data processing with foundation models, P2P‑T eliminates the need for human‑robot aligned data and achieves a 73% improvement over prior state‑of‑the‑art performance on complex real‑world tool manipulation tasks.
The paper investigates the numerical reliability of gradients in differentiable physics-based optimization for robotic material manipulation. Using two Material Point Method benchmarks, it identifies three key issues: GPU many-to-one sums that alter gradient signs, reduced reliability of finite-difference checks for long rollouts, and the impact of observation and loss definitions on optimization outcomes. The study recommends reproducible accumulation, finite-difference validation, and explicit objective reporting to improve robustness in robotic optimization.
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
InternW0-Δ is a unified World Action Model that integrates pretrained visual dynamics, scene semantics, 4D geometry, and motion priors within a Mixture-of-Transformers framework to generate robot actions. It leverages a frozen VLM for semantic guidance, a 4D foundation model for geometric priors, and introduces Causal Imprint to learn future-relevant scene changes without future-video rollout. The model is pretrained on a newly curated 20K‑hour heterogeneous corpus of robot and human demonstrations, achieving superior performance on simulation benchmarks and real‑robot platforms.
By Xingyu Miao, Zizun Li, Baole Fang, Kaiwen Song, Tenghui Wang, Hanxue Zhang, Yating Wang, Xudong Li, Yuping He, Xueyuan Wei, Chao Gao, Xijie Yang, Yingxiang Xu, Kerui Ren, Wenqi Guo, Jianjun Zhou, Xinzhe Wang, Weiguang Zhao, Ni Yang, Zetao Cai, Yufei Xue, Hengjie Li, Zeyu He, Yuanzhen Zhou, Rong Fu, Jianyang Zhang, Siwei Cui, Fuxian Huang, Yunsong Zhou, Xing Gao, Yifei Yao, Qiaojun Yu, Kailin Li, Ming Zhou, Mu Huang, Xinyue Li, Wenze Cui, Bingqi Jiang, Xueyue Zhu, Junting Dong, Haoyu Guo, Tao Lu, Mulin Yu, Bowen Zhou, Bin Zhao, Tianfan Xue, Weinan Zhang, Chunhua Shen
OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.
By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
The paper introduces an interaction‑centric framework that unifies representations for two‑finger gripper manipulation across different robot embodiments. By using a parameterized universal gripper abstraction and a canonical gripper‑frame representation, the system infers sub‑tasks from language and RGB‑D inputs, grounds interaction triplets, and employs hybrid features and a Flow‑Matching Transformer to generate smooth 7‑DoF action sequences. Experiments in both simulation and real‑world settings show that this approach achieves competitive benchmark performance while enabling extreme cross‑embodiment and cross‑viewpoint zero‑shot sim‑to‑real transfer to heterogeneous robot platforms.
By Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
By Nicol\'as Vera Z\'u\~niga
The paper introduces Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high‑resolution spatial representations from a YOLO11m‑P2 teacher to a lightweight YOLO11n student without changing the student’s inference architecture. CSCWD aligns teacher P2 features with student P3 while also applying same‑scale distillation at deeper pyramid levels, yielding a 2.92‑point mAP@0.5 improvement over the baseline and a 2.09‑point gain over same‑scale distillation alone. In zero‑shot tests on DUT‑Anti‑UAV and on a Raspberry Pi 5, the 2.58‑million‑parameter student reaches 50.32% mAP@0.5 at 82.32 ms latency (12.15 fps) with negligible runtime or memory increase.
By Amir Zamani, Zeinab Ghasemi-Naraghi
arXiv:2609.31452v1 Announce Type: cross
Abstract: Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and...
By Domen Tabernik, Peter Nimac, Jan Jeri\'cevi\'c, Danijel Sko\v{c}aj, Andrej Gams
arXiv:2504.18190v2 Announce Type: replace
Abstract: Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source d...
By Brun\'o B. Englert, Tommie Kerssies, Gijs Dubbelman