SafeFlow is a real‑time, text‑driven humanoid control framework that blends physics‑guided motion generation with a three‑stage safety gate. It uses Physics‑Guided Rectified Flow Matching in a VAE latent space to produce physically executable trajectories, accelerates sampling with Reflow, and filters unsafe outputs via semantic OOD detection, directional sensitivity checks, and hard kinematic constraints before handing them to a motion‑tracking controller. Experiments on the Unitree G1 show that SafeFlow achieves higher success rates, better physical compliance, and faster inference than diffusion‑ and retargeting‑based baselines while maintaining motion diversity.
By Hanbyel Cho, Sang-Hun Kim, Jeonguk Kang, Donghan Koo
MyoFlow introduces a discriminative flow-matching framework for high‑density surface electromyography (HD‑sEMG) gesture recognition that addresses distribution shifts caused by electrode re‑donning and physiological variability. By using a domain‑conditioned rectified flow to transport encoded windows toward gesture anchors, the method enables zero‑shot prediction without a separate classifier head. On the Hyser dataset, MyoFlow outperforms the strongest diffusion‑based baseline by 4.24 % in cross‑session accuracy and 6.37 % in cross‑subject accuracy, and achieves 91.71 % mean zero‑shot accuracy and 97.39 % mean few‑shot accuracy on the CEMHSEY dataset.
By Chenhao Wu, Dingjie Peng, Satoshi Funabashi, Satoshi Konishi, Wuqiang Yang, Hiroshi Onoda, Hironori Washizaki, Jiang Liu
The paper proposes a Friedkin‑Johnsen based framework to identify influential users in online social networks and assess how they shape community opinion. By manipulating initial opinions in experiments, the authors show that top influencers can significantly shift overall community sentiment, and their influence extends beyond direct neighbors to second‑degree contacts. The framework is validated on a tweet dataset from the U.S. presidential election, illustrating the power of digital influencers to alter public opinion.
By Omran Berjawi, Rida Khatoun, Giuseppe Fenza
The paper introduces the Neverwhere Visual Parkour Benchmark Suite, a collection of over sixty hyper‑photo‑realistic 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes designed to evaluate visual locomotion controllers in closed‑loop, continuous testing setups. It aims to bridge the gap between training and real‑world evaluation by providing reproducible environments and policy checkpoints trained across multiple scenes, while highlighting the risks of relying solely on 3D Gaussian‑generated data. The authors offer code and data on their project page for easy integration into robotic evaluation pipelines.
By Ziyu Chen, Henghui Bao, Haoran Chang, Alan Yu, Ran Choi, Kai McClennen, Gio Huh, Kevin Yang, Ri-Zhao Qiu, Yajvan Ravan, John J. Leonard, Xiaolong Wang, Phillip Isola, Ge Yang, Yue Wang
Intrinsic Robot Rewarding (IRR) leverages existing vision‑language‑action (VLA) systems to evaluate a robot’s own outcomes and provide feedback for policy improvement. By using successful demonstration endpoints as task‑specific references and the policy’s frozen visual encoder as the feature space, IRR adds a reference bank and scoring operation to the current pipeline without requiring a separate evaluator or additional perception backbone. The approach aims to reduce integration effort, reward computation cost, and human outcome scoring while enabling learning from the data already available in industrial robot systems.
By Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas, Mustafa Almohamad, Elham Al-Fuqara
The paper introduces ReImaGin, a method that uses image generation models as a flexible visual reasoning tool for multimodal large language models. Unlike traditional fixed-function vision tools, ReImaGin accepts natural language commands and can perform open-ended visual operations such as removing occlusions or creating floorplans from multiple views. Experiments on six diverse visual reasoning tasks show that ReImaGin outperforms both text-only reasoning and specialist vision-tool baselines, achieving up to a 25% improvement.
By Nishad Singhi, Hector Garcia Rodriguez, Aditya Arora, Marcus Rohrbach, Anna Rohrbach
MEgoVista is an offline pipeline that converts a single unprepared egocentric video into metric two‑hand and head motion within a gravity‑aligned world frame. It uniquely reconstructs motion in environments beyond studio volumes, uses calibrated stereo for absolute scale, and evaluates its outputs against independent optical capture to audit accuracy. The system thus expands the settings where high‑fidelity hand‑motion labels can be generated from natural, head‑worn recordings.
By Jiangong Xiao (Northwestern Polytechnical University), Zhihao Zhang (Xi'an Jiaotong University), Yifei Dong (Maniformer), Chao Ma (Maniformer), Zhouyi Jin (Maniformer), Zhiwen Hou (Maniformer), Li Liu (Maniformer), Weihuang Chen (Xi'an Jiaotong University), Hongbin Sun (Xi'an Jiaotong University), Maoqing Yao (Maniformer)
NeuroSymbEAD is a large‑scale neuro‑symbolic caption dataset that builds an ego‑centric knowledge graph of static and dynamic objects on the KITTI‑360 dataset, annotating classes, categories, heading directions, orientations, and distances from the ego‑vehicle. The dataset generates multilevel textual captions that serve as a lightweight representation of an ego‑centric scene map, enabling outdoor scene‑map reconstruction, visual recognition, and object grounding. Baselines for driving common sense and traffic/scene understanding are established, and the dataset is benchmarked using pre‑trained grounding and learned auto‑regressive captioning networks to support vision‑language and foundation models for traffic‑scene explanation, 3D reasoning, and interpretable autonomous‑driving perception.
By Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
GeoLAM is a framework that learns geometry‑grounded latent actions from unlabeled human videos. It uses future‑frame reconstruction with a frozen geometric feature hierarchy and motion supervision from a 4D geometry teacher to capture 3D displacement, image‑plane motion, and surface‑orientation changes. After pretraining, the representation serves as transition targets for a world‑action model trained on robot demonstrations, enabling denoised latent actions and executable action chunks without requiring hand‑pose annotations or future‑video generation during deployment.
By Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding
EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.
By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa
arXiv:2609.17387v1 Announce Type: new
Abstract: Real-time dense SLAM is a core capability for robotics applications that require robust localization and high- quality mapping in dynamic or fast-chang...
By Yongqi Mao, Hao Shi, Yufan Zhang, Zhonghua Yi, Xiangfei Guo, Kaiwei Wang
arXiv:2609.16319v1 Announce Type: cross
Abstract: Dexterous grasping is usually conducted for specific tasks, leading to heterogeneous constraints such as specific approach directions, desired contac...
By Hui Zhang, Mirko Meboldt, Jie Song
arXiv:2609.16686v1 Announce Type: cross
Abstract: Estimating deformable object states remains a fundamental challenge in robotics and simulation. We propose a novel factor graph-based framework for p...
By Lidia Al-Zogbi, Fangjie Li, Samuel Tobin, James Ferguson, Nithesh Kumar, Alejandro Chara, Kuan-I Chung, Mingxing Rao, Ayberk Acar, Susheela Sharma Stern, Robert Webster, Daniel Moyer, Alan Kuntz, Caleb Rucker, Tucker Hermans, Jie Ying Wu
arXiv:2609.16056v1 Announce Type: cross
Abstract: Humans carry behaviour knowledge of how to act in familiar situations into every new task rather than relearning it from scratch. There is no reason...
By Norbert Oswald, Fabian Deuser, Thomas Br\"aunl
arXiv:2609.16436v1 Announce Type: cross
Abstract: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the...
By Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan, Abhishek Nagaraj
arXiv:1906.07927v4 Announce Type: cross
Abstract: Deep neural networks (DNNs) have achieved great success in various applications due to their strong expressive power. However, recent studies have sh...
By Haonan Qiu, Chaowei Xiao, Lei Yang, Xinchen Yan, Honglak Lee, Bo Li
arXiv:2609.17141v1 Announce Type: cross
Abstract: Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain in...
By Hojin Lee, Yunho Lee, Daniel A Duecker, Cheolhyeon Kwon
arXiv:2609.16852v1 Announce Type: new
Abstract: Industrial IoT environments increasingly deploy autonomous mobile robots for tasks such as material handling, product assembly, or infrastructure inspe...
By Houssam Hajj Hassan (L2S), Antonia Maria Masucci (L2S), Lynda Zitoune (L2S), Salah-Eddine Elayoubi (L2S)
arXiv:2609.16075v1 Announce Type: cross
Abstract: Flexible robotic production requires joint decisions on process progression, material routing, resource assignment, temporary cooperation, and simult...
By Fouad Bahrpeyma, David Heik, Dirk Reichelt
arXiv:2609.16737v1 Announce Type: cross
Abstract: Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches of...
By Hojin Lee, Sizhe Lester Li, Maximilian Hilger, Susie Lu, Achim J. Lilienthal, Vincent Sitzmann, Daniel A. Duecker