X2-NativeCursor is a lightweight observer that tracks text progress in incremental‑text streaming TTS by aligning native speech tokens to the original text before waveform decoding. It uses a normalization plan and a local matcher to estimate the current label position, converting these estimates into a non‑backtracking cursor. The method achieves a mean absolute error of 0.151 Chinese characters with 80‑ms lookahead, outperforming a waveform‑based baseline and reducing alignment real‑time factor from 0.3598 to 0.0180.
By Zehan Liu, Carl Chen, Rime Wen, Kaiqi Fu, Altman Lin, Shawn Qin, Lights Shi, Roy Gan, Hao Wang, Qian Wang
arXiv:2609.07741v1 Announce Type: new
Abstract: Persistent AI assistants are intended to extend human attention, memory, and coordination across changing digital and physical environments. To be trul...
By Jo\~ao Dias Ferreira
arXiv:2609.07534v1 Announce Type: cross
Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained c...
By Yuhan Wang, Yurou Chen, Hongye Jiang, Wenzhao Lian
arXiv:2609.08123v1 Announce Type: cross
Abstract: A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly t...
By Suyog Khanal, Arun Kumar A V, Santu Rana
arXiv:2609.10506v1 Announce Type: cross
Abstract: Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However...
By Nisarga Nilavadi, Ralf R\"omer, Moritz Reuss, Michael Krawez, Tobias J\"ulg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
The paper presents HSAC, a reinforcement learning method that builds covering structures without predefined plans, using graph-structured states and a mixed action space of discrete block selection and continuous placement. It extends soft actor-critic with unilateral edges in graph neural networks to efficiently explore while simulating stability. Experiments show HSAC outperforms hybrid-PPO, remains robust to hyperparameters, and successfully transfers to a real two-robot 3D‑printed arch construction.
By Gabriel Vallat, Maryam Kamgarpour, Stefana Parascho
The paper introduces Reward Ensemble under Confidence (REC), a probabilistic reward learning framework for preference-based reinforcement learning that models per‑timestep reward uncertainty using an ensemble of distributional reward models. REC incorporates uncertainty into the preference loss and uses model disagreement to drive exploration, achieving 88.4% of shaped‑reward performance on acrobatic quadrotor control versus 55.2% with standard Preference PPO. The authors train policies in simulation and transfer them zero‑shot to real quadrotors, demonstrating complex acrobatic maneuvers learned solely from human preference feedback, and validate REC on a continuous‑control benchmark.
By Colin Merk, Ismail Geles, Jiaxu Xing, Angel Romero, Giorgia Ramponi, Davide Scaramuzza
arXiv:2508.05427v2 Announce Type: replace
Abstract: Large language models (LLMs) are beginning to reshape how organic-synthesis workflows are represented, queried, planned, and connected to experimen...
By Kartar Kumar, Rajesh Kumar, Nikesh Lagun
MamMA is a pedestrian trajectory prediction algorithm that leverages LiDAR-generated occupancy maps and egocentric vision sensor data. It partitions the occupancy map into patches to extract obstacle features and incorporates pedestrian awareness states, which influence perception and speed. Using a Mamba-based model, MamMA predicts future trajectories and outperforms state‑of‑the‑art methods on multiple benchmark datasets.
By Juncen Long, Xiaofeng Jin, Gianluca Bardaro, Simone Mentasti, Matteo Matteucci
arXiv:2609.09250v1 Announce Type: cross
Abstract: A verifier for robot policies reads a candidate behavior and returns a score for how well it did, used both to evaluate vision-language-action polici...
By Yang Wan, Xihang Yue, Zhirui Liu, Ziyuan Chu, Shuxun Wang, Yuhan Chen, Xiaonan Jiang, Xukun Zhu, Yubo Dong, Linchao Zhu
arXiv:2609.10377v1 Announce Type: cross
Abstract: Safety is a fundamental requirement for autonomous driving, yet existing end-to-end driving models still lack explicit risk-aware learning capacities...
By Yuanxin Tian, Zhiyuan Liu, Jinhao Li, Zhenhua Xu, Wenhao Yu, Jianqiang Wang
FALCON‑S is a modular, high‑fidelity simulator designed for fixed‑wing aerial robots operating near the ground. It models full 6DoF rigid‑body physics, semi‑empirical ground‑effect aerodynamics, actuator dynamics, sensor noise, and environmental disturbances, and supports both CPU and GPU backends via Torch and NVIDIA Warp for large‑scale reinforcement learning and optimal control. The framework offers a unified interface for various controllers, including RL and optical control algorithms, and allows cross‑validation with X‑Plane and JSBSim for engineering integration and visual fidelity.
By Matteo El Hariry, Pedro Lima, Andrej Orsula, Matthieu Geist, Miguel Olivares-Mendez
arXiv:2605.03927v3 Announce Type: replace
Abstract: Vision-language models have demonstrated strong performance across robotic perception and instruction-following tasks. However, they still struggle...
By Xiaowen Sun, Matthias Kerzel, Mengdi Li, Xufeng Zhao, Paul Striker, Stefan Wermter
arXiv:2609.08800v1 Announce Type: cross
Abstract: Three properties determine whether a differentiable simulator can drive gradient-based optimization through contact: simulation accuracy, gradient re...
By Ale\v{s} Ku\v{c}era, Karel Zimmermann
arXiv:2609.08602v1 Announce Type: new
Abstract: Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequenc...
By Tianyi Ma, Parisa Kordjamshidi
arXiv:2609.09137v1 Announce Type: new
Abstract: Robotic Process Automation (RPA) is widely used to reduce administrative burden in United States hospitals, yet an estimated 30-50% of RPA initiatives...
By Maria Alejandra Gomez, Juan Manuel Castillo
The paper introduces Distributed Dexterous Manipulation (DDM), a challenging control problem involving 64 soft delta robots arranged in an 8x8 grid. It presents a framework using spatially conditioned Multi-Agent Transformers (MATs) with adaptive layer norm, spatial contrastive embeddings, and a behavior cloning method fine‑tuned by Soft Actor Critic. Experiments demonstrate that MATs refine actions through stacked attention blocks, enabling long‑horizon planar manipulation in simulation and real‑world settings, while an action‑selection strategy reduces robot usage by about 65% and lowers wear‑and‑tear, achieving an average error of ~1.5 cm.
By Sarvesh Patil
TimeBlind is a diagnostic benchmark designed to evaluate fine‑grained spatio‑temporal compositionality in video large language models (LLMs). It categorizes temporal understanding into three levels—atomic event recognition, event property characterization, and reasoning about event interdependencies—and uses a minimal‑pairs paradigm where video pairs share identical static content but differ only in temporal structure. Across 20 state‑of‑the‑art MLLMs tested on 600 curated instances, the best model achieved only 48.2% instance accuracy, far below human performance of 98.2%, highlighting a reliance on static visual shortcuts rather than true temporal reasoning.
By Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius
arXiv:2609.05593v1 Announce Type: cross
Abstract: Autonomous robot navigation failures differ not only in categorical severity but also in the physical context in which they occur. A near-miss at low...
By Rifa Ferzana
arXiv:2609.05533v1 Announce Type: cross
Abstract: Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earl...
By Cheng Yin, Wang Xu, Junpeng Yang, Sikyuen Tam, Hanyu Liu, Yuan Yao, Xiangrui Zeng, Junbo Cui, Yequan Wang, Zhouping Yin, Yankai Lin