DGSG-Mind introduces a hybrid instance-aware 3D Gaussian dynamic scene graph system that integrates open‑vocabulary semantic information into dynamic 3D scene representations. By coupling a probabilistic voxel grid with explicit 3D Gaussians, it achieves robust cross‑modal instance fusion, incremental semantic mapping, and dynamic change handling through Gaussian‑based relocalization and masked refinement. The system builds a hierarchical scene graph and a 3D Gaussian Mind for multimodal reasoning, achieving state‑of‑the‑art zero‑shot 3D visual grounding and strong performance in open‑vocabulary semantic segmentation and scene reconstruction, and is demonstrated on real‑world robots.
By Luzhou Ge, Xiangyu Zhu, Jinyan Liu, Xuesong Li
The paper introduces LCAP, a method for adapting photonic neural networks to real hardware by learning a shared correction from a population of chips and then personalizing each chip using only 32 fixed output probes. LCAP decomposes adaptation into a transferable population correction and a probe‑inferred latent personalization, allowing feed‑forward calibration without device‑specific optimization. Experiments on a simulated three‑layer 64‑mode MZI network show accuracy improvements from 80.4% to 93.4% and significant gains on unseen chips.
By Tianyu Gao, Guantian Zheng
Auto-HSI is a system that creates personalized human‑swarm interaction interfaces on demand using large language models to automatically generate code from natural language descriptions and gesture demonstrations. The prototype enables untrained operators to control a swarm of 50 simulated robots with one‑ or two‑hand gestures, allowing teleoperation of motion, formation shape, and shape deformation. Experiments demonstrate the system’s gesture tracking, code generation, and live operation capabilities, including real‑time updates and deployment on real robots.
By Alessandro Nazzari, Nathan Cerisara, Dorian Tonnis, Raina Zakir, Lorenzo Labarile, Weixu Zhu, Marco Dorigo, Mary Katherine Heinrich
ProxiDex is a dynamics‑guided proximity policy framework for multi‑finger dexterous manipulation that treats hand‑object proximity as an interaction state. It reconstructs interaction point clouds, converts geometric distances into proximity cues, and learns action‑conditioned proximity dynamics using a coupled forward‑inverse design. The framework adaptively reweights proximity tokens across manipulation phases and employs dynamics‑consistency supervision to stabilize action generation, leading to improved success rates and robustness in both simulation and real‑world experiments.
By Yushan Bai, Boyu Zheng, Zhiyang Mao, Hongzheng Sun, Yuchuang Tong, En Li, Zhengtao Zhang
arXiv:2609.16637v1 Announce Type: cross
Abstract: Efficient perception is central to robotic systems operating under constrained computation, memory, and latency budgets. Knowledge transfer from larg...
By Yanick C. Tchenko, Felix Mohr, Hicham Hadj-Abdelkader, Hedi Tabia
arXiv:2609.16683v1 Announce Type: cross
Abstract: Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and objec...
By Liu Cao, Xingze Wu, Jingzhi Cui, Botian Xu, Mingzhi Pei, Ruoqu Chen, Mengdi Xu
arXiv:2609.16369v1 Announce Type: new
Abstract: Precise manipulation of liquid droplets underpins lab-on-a-chip platforms for diagnostics, chemical synthesis, and biological assays. Yet autonomous dr...
By Rajneesh Anand, Mayuresh V. Kothare
arXiv:2609.16145v1 Announce Type: new
Abstract: We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? W...
By Gautam Kishore
arXiv:2609.17419v1 Announce Type: new
Abstract: Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We study the...
By Xinyuan Song, Zekun Cai
arXiv:2609.16697v1 Announce Type: cross
Abstract: World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interv...
By Nanjie Yao, Hao Wang, Chong Cheng, Zhikang Chen, Wenzhe Li, Jiafei Lyu, Li Shen, Peilin Zhao, Zongqing Lu, Gao Huang, Steven Hoi, Dacheng Tao, Deheng Ye
arXiv:2609.16186v1 Announce Type: cross
Abstract: Autonomous soft-tissue cancer surgery has been limited to interventions on organ surfaces, because current systems cannot perceive and adapt to anato...
By Ethan Kilmer, Pit Henrich, Jiawei Ge, Paul M. Scheikl, Laura Connolly, Soum D. Lokeshwar, Joseph Chen, Justin D. Opfermann, Kaitlyn Kumar, Lauren Shepard, Ahmed Ghazi, Nirmish Singla, Richard J. Cha, Kevin Cleary, Franziska Mathis-Ullrich, Axel Krieger
The paper introduces QDTraj, a method that uses Quality‑Diversity algorithms to automatically generate a diverse set of low‑level trajectory primitives for manipulating articulated objects. By leveraging sparse reward exploration, QDTraj produces at least five times more diverse trajectories for hinge and slider tasks compared to baseline methods, and demonstrates strong generalization across 30 articulations from the PartNetMobility dataset, averaging 704 trajectories per task. The resulting primitives are validated both in simulation and on real robots, with the code released publicly.
By Mathilde Kappel, Mahdi Khoramshahi, Louis Annabi, Faiz Ben Amar, St\'ephane Doncieux
HumanEgo is a framework that enables zero‑shot robot learning from short egocentric human videos by converting each demonstration into an entity‑level hand‑object interaction representation and training a flow‑matching policy with dense auxiliary objectives. The method is robot‑data‑free, hardware‑agnostic, and data‑efficient, achieving 92.5 % success on four real‑world tasks with only 30 minutes of human video per task and outperforming matched‑time robot teleoperation by 41 %. HumanEgo also robustly transfers zero‑shot across new robots, cameras, and environments, and is released as an open‑source tool for learning robot policies directly from human data.
By Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos
The paper introduces a new paradigm called VLM-as-probabilistic-grounder, which models the uncertainty of Vision‑Language Model (VLM) predicate groundings as a probability distribution over symbolic states. This probabilistic grounding allows belief‑space planning, producing more robust plans in partially observable settings. Experiments in simulated household robot environments demonstrate that this approach improves robustness and task success compared to deterministic grounding methods.
By Guy Azran, Michael Navat, Sarah Keren
The article argues that Surprisal Theory, often presented as a computational-level explanation, is not a theory in its own right. It contends that using large language model (LLM) surprisals without considering the underlying representational and algorithmic choices obscures the theory’s commitments. The authors demonstrate through three analyses that algorithm and architecture significantly influence language model probabilities, urging researchers to reassess treating LLM surprisals as interchangeable.
By Andr\'es Bux\'o-Lugo, Aniello De Santo, Morgan Grobol, Ryan J. Hubbard, Cassandra L. Jacobs
EgoPathBench is a new dataset and benchmark that tests zero‑shot egocentric waypoint decision‑making in vision‑language models. Each task presents an egocentric RGB image, a natural‑language goal, and numbered visible waypoints, and models must return traversable candidates or an ordered route. The benchmark evaluates candidate feasibility, edge legality, and goal arrival under point‑agent or embodied geometry, covering 31,852 training, 1,345 validation, and 1,111 benchmark questions.
"whyItMatters":"The benchmark reveals that current VLMs perform poorly on integrated navigation tasks, highlighting a gap in spatial intelligence that can be addressed by fine‑tuning with the released training data."
By Yang Zhao, Zhuo Chen, Xubo Yang
FluxVLA Engine is an open, configuration‑driven platform that unifies the fragmented components of embodied policy development—datasets, visual‑language and world models, action heads, learning methods, distributed training, simulation evaluation, inference, and robot interfaces—into a reproducible data‑to‑deployment workflow. It adds features such as compositional dual‑arm simulation, scalable automatic data generation, human‑in‑the‑loop rollout and correction, Real‑Time Chunking for fast inference, and lightweight remote GPU serving, thereby linking offline learning, simulation validation, online correction, and real‑robot execution under shared, auditable contracts. The engine aims to eliminate engineering bottlenecks that currently separate promising embodied‑learning algorithms from reliable, reproducible deployment.
By Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang, Chenguang Cui, Hongyuan Chen, Xu Huang, Zheyi Zhao, Pinxi Shen, Bozhen He, Zhen Fu, Yifan Wang, Zexin Zhang, Ang Gao, Haoyu Chen, Chengqi Shi, Hua Chen
arXiv:2609.17521v1 Announce Type: cross
Abstract: Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet...
By Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu
The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.
By Xiaoyan Li, Yunli Wang
The paper introduces CTAN, a Cycle-Temporal Attention Network for audio‑visual embodied navigation. It proposes an Audio‑Visual Reconstruction Cross‑Attention module that uses bidirectional cycle‑consistency to strengthen spatial semantics across visual and acoustic modalities, and a Temporal Cross‑Modal Memory to fuse real‑time multimodal features with historical context. Experiments on Replica and Matterport3D show that CTAN outperforms prior methods in success rate, SPL, and scene navigation accuracy.
By Teng Liu, Yinfeng Yu