DeCAL is a vision‑language‑action model designed for dexterous manipulation that incorporates tactile sensing through adaptive visuo‑tactile fusion and latent co‑imagination. It uses a Mixture‑of‑Transformers architecture with specialized experts for understanding, imagination, and action, enabling efficient information flow and dynamic regulation of tactile inputs. Experiments show DeCAL achieves state‑of‑the‑art performance, with a 71% average success rate and 83.4% progress success rate, and generalizes well to unseen scenarios.
DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.
The paper presents a reinforcement learning method, HSAC, that builds covering structures without relying on rigid, pre‑planned sequences. It uses graph‑structured state representations and a mixed action space to select blocks and adjust their placement continuously, while an efficient exploration strategy incorporates unilateral edges into graph neural networks. HSAC outperforms the prior hybrid‑PPO approach, shows strong sample efficiency, robustness to hyperparameters, and successfully transfers policies from simulation to a real two‑robot 3D‑printed block construction task.
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executab...
ReMoMask-2 is a retrieval‑augmented text‑to‑motion generation framework that improves on complex motion descriptions by addressing coarse retrieval and representation gaps. It introduces a structure‑aware RAG pipeline with Hierarchical Bidirectional Momentum contrastive learning, Semantic Spatial‑Temporal Attention, and Topology Structured Masking, and rebuilds the retrieval database in the generator’s latent space using a lightweight projector. Experiments on HumanML3D, KIT‑ML, and SnapMoGen show state‑of‑the‑art retrieval accuracy and the lowest FID scores, with a single mask‑transformer stage delivering faster inference than the previous two‑stage design.
FPicker is a topology‑guided framework for filament tracing in low‑signal Cryo‑EM images. It combines a center‑endpoint representation with an open‑curve evolution module to model non‑cyclic connectivity, overcoming limitations of pixel‑wise segmenters, box‑based detectors, sequential trackers, and traditional active contours. On simulated benchmarks, FPicker improves mean spatio‑angular precision by over 40 % and reduces topological gap rates by more than 60 % under extreme noise, and it achieves state‑of‑the‑art performance on real EMPIAR data after fine‑tuning.
The paper introduces 3DWay, a method that predicts 3D consistent waypoints for robot manipulation using multi‑view images. By first generating 2D waypoints that are consistent across views and then triangulating them, the approach provides explicit 3D motion specifications while leveraging pretrained vision‑language models. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing the generalization of robot manipulation policies.
We present our solution to the LUMPI track of the UCF UrbanTwin Sim2Real LiDAR Challenge at the 6th DriveX Workshop, ECCV 2026. The detector must be trained only on synthetic data and is evaluated on...
This paper presents a novel framework designed to enhance key object identification in autonomous driving. Existing methods primarily focus on either detecting objects independently or leveraging visu...
The paper proposes a neuro‑symbolic framework that augments vision‑language‑action (VLA) models with explicit task graphs and multimodal procedural memory to handle long‑horizon manipulation tasks. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory tracks the active step, completed actions, textual context, and relevant visual evidence. Human demonstrations provide spatial and temporal guidance via gaze or saliency cues, which are annotated in robot‑view teleoperation videos and used to fine‑tune VLA models. The approach is evaluated on workspace clearing and surgical‑instrument handling tasks, measuring object and destination selection, subtask completion, task progress, step‑order consistency, overall success, and procedural or execution mistakes.
By Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, J\"org Kr\"uger
Squint is a visual Soft Actor Critic algorithm designed to accelerate reinforcement learning for robotics. It combines parallel simulation, a distributional critic, resolution squinting, layer normalization, a tuned update-to-data ratio, and an optimized implementation to reduce wall‑clock training time. On the SO‑101 Task Set, Squint trains policies in as little as 15 minutes on a single RTX 3090 GPU, with most tasks converging in under 6 minutes and successfully transferring to a real SO‑101 robot.
By Abdulaziz Almuzairee, Henrik I. Christensen
The paper introduces E‑RGB‑D, a real‑time event‑based perception system that combines a Digital Light Processing projector with a monochrome event camera to produce RGB‑D data. By projecting structured light and capturing asynchronous brightness changes, the system can detect color and depth for each pixel, achieving a color detection speed of 1400 fps and a depth detection rate of 4 kHz. The approach enables frameless RGB‑D sensing and delivers colorful point clouds without compromising spatial resolution.
By Seyed Ehsan Marjani Bajestani, Giovanni Beltrame
The paper reports the first empirical study comparing how humans and large language models (LLMs) evaluate perceived moral agency (PMA) in both human and autonomous artificial agents within smart city scenarios. Using a validated PMA scale, 190 human participants and various LLMs were assessed, revealing that humans are perceived to have higher moral agency than artificial agents. When confronted with moral dilemmas, LLMs focus on situational factors such as harm severity and urgency, mirroring the context‑sensitivity observed in human raters.
By Fernanda Mansilla, Aloysius Tok, Bahia Guella\"i, Farah Benamara, Nancy F. Chen
VoxelFix is a graph‑based post‑hoc semantic correction method that refines voxel labels in completed 3D voxel maps while preserving their geometry and occupancy. It learns to correct errors by exploiting local geometry and neighboring semantic information, using training pairs generated by corrupting annotated maps with class confusions from upstream perception pipelines. Experiments on OccuFly maps show consistent improvements of 4.23–5.00 percentage points in mIoU, especially for tree, roof, and wall classes, and the method generalizes to out‑of‑distribution aerial scenes.
By Sunesh Praveen Raja Sundarasami, Taehyoung Kim, Johannes Scherer, Toma\v{z} Coti\v{c}, Sivasubiramaniam Subbiah, Andreas Greiner, Paul Spannaus, Sebastian Houben
The paper explores the link between understanding and compression, arguing that while compression is a useful proxy for robust competence, it is not identical to comprehension. It proposes that true understanding involves a mental model of relational structure that enables prediction, and that this predictive ability underlies compression. The authors further suggest that human understanding is shaped by the need for demonstrability and transmissibility, leading to a preference for principled simplicity.
By Matthieu Queloz, Pierre Beckmann
The paper introduces a resilient remote robotic control framework that couples control and wireless world models using a Joint Embedding Predictive Architecture (JEPA). By learning latent representations from visual observations and radio frequency (RF) data, the system predicts future robot states and wireless conditions to schedule uplink transmissions efficiently. An adaptive resilience mechanism adjusts perception embeddings when prediction errors arise, enabling robust operation without retraining the entire control policy.
By H. P. Madushanka, Sumudu Samarakoon, Mehdi Bennis
The paper addresses performance discrepancy in cross-domain 3D class‑incremental learning, where 3D point clouds from heterogeneous sources cause varying degrees of performance loss beyond catastrophic forgetting. The authors introduce the Domain3D‑CIL protocol and adapt existing CIL methods to 3D, showing consistent discrepancy across baselines. They propose PolyMem, an exemplar‑free approach that models high‑order feature statistics to improve cross‑domain robustness and reduce performance discrepancy.
By Jinge Ma, Gautham Vinod, Bruce Coburn, Jui-Feng Chi, Siddeshwar Raghavan, Fengqing Zhu
The paper examines whether large language models (LLMs) exhibit conspiratorial tendencies, socio-demographic biases in this domain, and how easily they can be conditioned to adopt conspiratorial viewpoints. Using validated psychometric surveys, the authors find that LLMs partially align with conspiracy beliefs, that conditioning with demographic attributes yields uneven effects revealing latent biases, and that targeted prompts can readily shift responses toward conspiratorial stances. These findings underscore the vulnerability of LLMs to manipulation and the potential risks of deploying them in sensitive contexts.
By Francesco Corso, Francesco Pierri, Gianmarco De Francisci Morales
RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.
By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
The paper introduces Linguistic Trajectory Encoding (LTE), a hybrid representation that compresses dynamic object motion histories using natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location while preserving accuracy with geometric waypoints and linguistic descriptions. Evaluated on the newly constructed Spatial Memory Benchmark (SMB) from EgoLife multi‑day recordings, LTE achieves 45.3 % success in semantic trajectory retrieval and 48.7 % in long‑horizon object retrieval, outperforming prior structured‑memory and VLM baselines, and compresses trajectories 8.7×–26.1× with sub‑second query latency on 24‑hour video.
By Tianyidan Xie, Shenyi Wang, Qiang Tang, Mingjie Wang, Zhicheng Qiu, Xuanfu Li, Zhan Xu, Jian Yang, Lanjun Wang, Zili Yi