arXiv:2609.17628v1 Announce Type: cross
Abstract: Active perception allows autonomous agents to select their viewpoints rather than passively process the viewpoints given to them, enabling them to ta...
By \"U. Bora G\"okbakan (WILLOW), St\'ephane Caron (ISIR), Philippe Sou\`eres (LAAS-GEPETTO)
Changepoint-Aware World Models (CAWM) is a DreamerV3 agent that detects abrupt dynamics shifts in a robot’s environment using an online CUSUM test on internal prediction error. Upon detection, CAWM selectively forgets stale replay data while preserving the learned representation, enabling rapid recovery from shifts such as doubled gravity or halved actuator gain. Experiments on simulated locomotion show CAWM recovers faster than passive retraining and outperforms a baseline that respawns a fresh dynamics model, achieving significant return gains in the first 30k post‑shift frames.
By Everest Yang
The paper introduces CSWAM, a Causal Semantic World Action Model that enhances FastWAM by integrating a causal semantic expert based on V-JEPA 2.1. This expert provides temporally grounded, appearance‑agnostic representations of semantic state changes and motion, leveraging sparse observation history and causal attention to improve action‑only inference. Experiments on simulation and real‑robot tasks show that CSWAM significantly boosts out‑of‑distribution generalization, raising success rates from 10.16% to 45.18% on RoboTwin 2.0 and from 27.5% to 70.0% across real‑robot tasks.
By Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu
arXiv:2608.13014v2 Announce Type: replace
Abstract: Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet...
By Andela Ilic, Rachel Schuchert, Yijing Jiang, Christian Holz
The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.
By Andrew P. Berg, Qian Zhang, Mia Y. Wang
We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Spla...
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into do...
Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information,...
Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies...
Self-driving laboratories can explore synthesis conditions autonomously, but their decision-making layer is typically a black-box optimizer, and the output is a set of optimized samples, with the meas...
Collective intelligence is a collaborative autonomy paradigm in which multiple agents pursue shared objectives through local perception, information exchange, and coordinated action. UAV swarms embody...
Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, yet often leave metric relations among actions, objects, and scene geometry implicit. Hum...
This paper studies energy-aware manipulation as a physically grounded learning problem. We define a joint-space mechanical-work proxy from joint torque and angular displacement, and train a differenti...
Cross-view object geo-localization (CVOGL) locates a target in satellite imagery using drone or street-view queries. Existing methods train separate detectors for each viewpoint, leading to parameter...
The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI). These agentic AI...
HumanEgo is a framework that enables zero‑shot robot learning from short egocentric human videos by converting each demonstration into an entity‑level hand‑object interaction representation and training a flow‑matching policy with dense auxiliary objectives. The method is robot‑data‑free, hardware‑agnostic, and data‑efficient, achieving 92.5 % success on four real‑world tasks with only 30 minutes of human video per task and outperforming matched‑time robot teleoperation by 41 %. HumanEgo also robustly transfers zero‑shot across new robots, cameras, and environments, and is released as an open‑source tool for learning robot policies directly from human data.
By Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos
The paper introduces QDTraj, a method that uses Quality‑Diversity algorithms to automatically generate a diverse set of low‑level trajectory primitives for manipulating articulated objects. By leveraging sparse reward exploration, QDTraj produces at least five times more diverse trajectories for hinge and slider tasks compared to baseline methods, and demonstrates strong generalization across 30 articulations from the PartNetMobility dataset, averaging 704 trajectories per task. The resulting primitives are validated both in simulation and on real robots, with the code released publicly.
By Mathilde Kappel, Mahdi Khoramshahi, Louis Annabi, Faiz Ben Amar, St\'ephane Doncieux
ProxiDex is a dynamics‑guided proximity policy framework for multi‑finger dexterous manipulation that treats hand‑object proximity as an interaction state. It reconstructs interaction point clouds, converts geometric distances into proximity cues, and learns action‑conditioned proximity dynamics using a coupled forward‑inverse design. The framework adaptively reweights proximity tokens across manipulation phases and employs dynamics‑consistency supervision to stabilize action generation, leading to improved success rates and robustness in both simulation and real‑world experiments.
By Yushan Bai, Boyu Zheng, Zhiyang Mao, Hongzheng Sun, Yuchuang Tong, En Li, Zhengtao Zhang
The paper introduces a new paradigm called VLM-as-probabilistic-grounder, which models the uncertainty of Vision‑Language Model (VLM) predicate groundings as a probability distribution over symbolic states. This probabilistic grounding allows belief‑space planning, producing more robust plans in partially observable settings. Experiments in simulated household robot environments demonstrate that this approach improves robustness and task success compared to deterministic grounding methods.
By Guy Azran, Michael Navat, Sarah Keren
arXiv:2609.16319v1 Announce Type: cross
Abstract: Dexterous grasping is usually conducted for specific tasks, leading to heterogeneous constraints such as specific approach directions, desired contac...
By Hui Zhang, Mirko Meboldt, Jie Song