REARL is a closed‑loop simulation enhancement framework that combines real traffic data with large language models (LLMs) to improve autonomous driving simulations. It clusters real traffic, uses cluster centers as representative scenarios for the LLM, and employs a sliding‑window detector to monitor vehicle speed and spacing discrepancies. When thresholds are exceeded, the LLM adjusts vehicle decision‑making or selects matching real vehicle actions, resulting in lower Hellinger distance and MAPE compared to baselines in a HighD highway setting.
By Xiaojun Bi (Minzu University of China, Beijing, China), Jun Jiang (Minzu University of China, Beijing, China), Yiwen Sun (Peking University, Beijing, China, BIGAI, Beijing, China), Quanyi Ou (Minzu University of China, Beijing, China), Ke Cheng (Beihang University, Beijing, China), Mingjie Bi (BIGAI, Beijing, China), Yexin Li (BIGAI, Beijing, China)
GLAMDRING is a framework that jointly designs a quadruped robot’s morphology and its gait controller using reinforcement learning of Hopf-oscillator Central Pattern Generators (CPGs). Given specifications such as forward‑velocity bounds, actuator power budgets, an actuator library, and payload requirements, the system returns an optimized robot design and a corresponding gait policy, ranking designs by objectives like maximum speed, minimum Cost of Transport, or maximum payload margin. Experiments demonstrate that co‑designing body and gait is essential for meeting locomotion constraints, that actuator feasibility determines payload capacity, and that natural animal gaits emerge from the design process, with a real‑world demonstration confirming the approach’s effectiveness.
By Amogh Joshi, Kaushik Roy
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
By Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang
SCOUT is a frozen‑encoder approach for sim‑to‑real text‑based person retrieval that predicts cross‑modal embeddings instead of fine‑tuning cross‑encoders. It uses a trainable predictor to map patch tokens from a frozen video encoder (V‑JEPA) into the embedding space of a frozen text encoder (EmbeddingGemma), guided by a bidirectional InfoNCE objective. The method achieves state‑of‑the‑art results on the AI City Challenge 2026 Track 4, with a full retrieve‑fuse‑rerank pipeline reaching 84.25 mAP@10 and a single frozen model alone scoring 60.63, while training costs are modest (≈95 GPU‑hours).
By Abdarahmane Traor\'e, Andy Couturier, \'Eric Hervet
MM-Future is a world-action model for autonomous driving that generates multiple paired scene-action hypotheses and captures bidirectional interaction within each pair. It initializes each hypothesis from a structured action prior and an independent future scene source, then co-evolves them using a modality-aware diffusion Transformer. The model compresses multi-view video into planning-oriented MM-Tokens and uses a future-conditioned proposal scorer to rank trajectory candidates, achieving strong performance on NAVSIM and HUGSIM benchmarks.
By Shuai Liu, Hechangle Gong, Hao Jiang, Runlin He, Junxiang Zhan, Kai Huang, Sheng Yang, Shaoqing Ren
CoRef-GS introduces a cooperative referring Gaussian splatting framework for multi‑agent scene understanding, enabling robots to ground object‑ and relation‑centric language queries across independently reconstructed maps. The method builds local open‑vocabulary instance‑aware Gaussian maps, aligns them using a cross‑agent module that enforces geometric and semantic consistency, and grounds queries with a view‑conditioned mask relation graph. Experiments on a new dual‑quadruped benchmark show significant improvements, reducing rotation error from 2.58° to 0.15° and raising real‑world referring mIoU from 52.6% to 68.8%.
By Zhikun Zhou, Kunyu Peng, Runyi Yang, Junhao Cai, Di Wen, Ruiping Liu, Danda Pani Paudel, Yi Zhou, Luc Van Gool, Kailun Yang
The paper introduces an extension to the OceanSim underwater perception simulator, adding a Synthetic Data Generation pipeline that produces large, automatically labeled, photorealistic datasets with configurable scene and sensor settings. The authors evaluate this pipeline on a real-world sea urchin detection task, examining how different synthetic scene variations influence sim-to-real performance. They discuss the pipeline’s findings, limitations, and future directions for improving rendering fidelity, scene diversity, and sim-to-real generalization.
By Haoyu Ma, Onur Bagoren, Anja Sheppard, Elias Fandi, Ashrith Edukulla, Tanner Aslan, Natasha Sieh, Jingyu Song, Katherine A. Skinner
Astronex-World 1.0 is an open, controllable video world‑model foundation that can generate future visual states from text prompts or initial images, supporting frame‑aligned camera trajectories, continuous actions, and text events during rollout. It offers both a bidirectional model for full‑context generation and a causal model with block‑causal attention, built on the Wan2.2‑TI2V‑5B prior, and achieves real‑time 832×480 video at 24 fps. The model was trained over five stages on two NVIDIA L20 GPUs, scores 73.5 on WBench Navi and 70.0 on WBench Full, and outperforms several larger competitors while providing interfaces for embodied intelligence and autonomous driving.
By Xin Zhou, Cong Miao
UniExo is a framework that builds a single, multi-skill musculoskeletal human policy by distilling four imitation experts—walking, turning, running, and backward walking—into one network guided by a skill latent. The human policy is fine‑tuned with reinforcement learning on transition sequences, achieving a 94.7% tracking success rate on unseen clips and greater robustness to perturbations. A hip exoskeleton controller is then co‑adapted with this human policy via multi‑agent reinforcement learning, enabling it to assist across four treadmill speeds and a continuous route of all four skills without explicit mode switching.
By Yifei Yuan, Jakob Wolf, Ghaith Androwis, Xianlian Zhou
The paper presents a deep‑learning perception framework for selective robotic cotton harvesting, evaluated on 1,008 field images captured under diverse lighting and weather conditions. Detection models from YOLOv8 to YOLOv13 were benchmarked, with GELAN‑s achieving the best trade‑off between accuracy and speed. For segmentation, YOLOv12‑m‑seg outperformed other models, and a detection‑prompted segmentation approach using GELAN‑s bounding boxes further improved localization for SAM variants. Field trials with a UR5e robot and ZED2i camera confirmed YOLOv12‑m‑seg’s real‑time performance for cotton boll detection, segmentation, and selective picking.
By Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins
FunArt is a framework that builds articulation‑aware functional 3D scene graphs from a single static RGB‑D observation. It reconstructs object instances, converts their geometry into the O‑Voxel representation of TRELLIS.2, and uses a frozen sparse‑compression VAE as a structural prior. A lightweight query‑based decoder jointly segments movable parts and functional interactive elements while estimating motion type, axis, origin, and range, achieving state‑of‑the‑art results on the Articulate3D dataset.
By Dennis Rotondi, Abdelrhman Werby, Kai O. Arras
The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.
By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu
The paper introduces DynaWeightPnP, a real‑time algorithm for correspondence‑free Perspective‑n‑Point (PnP) problems that aligns 3D and 2D shapes without needing point correspondences. It uses a Reproducing Kernel Hilbert Space formulation solved via iterative reweighted least squares, and addresses a numerical ambiguity between rotation and translation with a dynamic weighting sub‑problem and alternative search strategy. Experiments on 3D‑2D vascular centerline registration in endovascular image‑guided interventions show processing rates of 60 Hz (without refinement) and 31 Hz (with refinement) on a single‑core CPU, achieving accuracy comparable to existing methods.
By Jingwei Song, Maani Ghaffari
SceneTeract is a verification interface that separates semantic action understanding from physical feasibility in indoor 3D scenes. It decomposes activities into atomic actions and performs explicit geometric checks to determine executability, providing diagnostic traces for failures. The system reveals widespread functional and accessibility issues in synthetic scenes, shows that existing VLMs over‑predict action feasibility, and improves VLM performance through post‑training with verifier feedback, with benefits that generalize to real‑world scenes.
By L\'eopold Maillard, Francis Engelmann, Tom Durand, Boxiao Pan, Yang You, Leonidas Guibas, Maks Ovsjanikov
The paper introduces a privacy‑preserving approach to open‑vocabulary 3D semantic segmentation that operates solely on depth data, eliminating the use of RGB images to avoid disclosing scene‑specific visual information. It proposes a stricter depth‑only evaluation protocol and presents UTTO, a model‑agnostic uncertainty‑guided test‑time optimization framework that refines predictions from frozen open‑vocabulary 3D backbones using structured predictive uncertainty. Experiments on ScanNet and Matterport3D show consistent improvements, and additional analyses demonstrate the method’s relevance for privacy‑constrained robotic applications.
By Xuying Huang, Sicong Pan, Maren Bennewitz
CitySTAR introduces a training‑free framework that transforms billion‑scale urban point clouds into a query‑ready scene graph of open‑vocabulary 3D instances, using CodeLLM‑driven tools to supply multimodal evidence for node attributes and spatial relations. It models target‑context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation, followed by a Reflective Cross‑modal Grounding module that integrates topology consistency and 2D visual evidence to decide over a metric‑aware 3D context graph. The authors also present CitySTAR‑3D, a benchmark that enhances semantic coverage, instance completeness, bounding‑box fidelity, and spatial‑relation complexity for city‑scale 3D grounding, and report extensive experiments showing consistent improvements in open‑world urban 3D grounding with strong interpretability and generalization.
By Shuai Zhang, Hongye Hou, Qinghe Liu, Zhuoxiao Li, Dongli Wu, Jing Ou, Yuan Liu, Wufan Zhao
GAVEL is a framework that uses an explicit graph world model to verify and repair long‑horizon plans generated by large language models (LLMs). The graph encodes object relations, action pre‑conditions and effects, and probabilistic beliefs about unobserved object locations, allowing the system to predict action outcomes, detect violations, and repair them before execution. In experiments on BEHAVIOR‑1K, GAVEL boosts single‑task success from 41.2 % to 91.8 % and multi‑task success from 19.9 % to 92.6 %, while also reducing travel distance by about 5.4 % compared with a static variant.
By Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic
CoreSense is a robot‑system integration architecture that traces episodic evidence and uses a conflict‑aware belief gate to decide whether to proceed, re‑observe, abstain, or escalates. The gate evaluates scope, provenance, time, contradiction, and support before making a recommendation. Evaluation on public robot datasets, simulations, and a live cloud deployment shows that belief gating can eliminate protocol‑defined unsafe proceeds while maintaining auditability.
By Zoe Li
The paper introduces Movement Trend Guidance, a method that equips 3D diffusion policies with foresight by learning a compact latent representation of interaction evolution from a brief observation history. This latent, supervised by sparse future gripper states during training, serves as future-oriented conditioning during inference, enhancing action generation without adding explicit planning. The approach improves performance on RoboTwin2.0, LIBERO-40, and DexArt benchmarks, achieving higher success rates across multiple tasks.
By Zhongbo Zhang, Zaibin Zhang, Yifan Wang, Changbo Yan, Lijun Wang, Huchuan Lu
The paper introduces a lightweight full-frame detector for partially manipulated AI-generated videos, suitable for edge deployment without face-detection preprocessing. It distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student using temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter. The model addresses false positives on legitimate scene cuts and threshold-level miscalibration, achieving an AUC of 0.766 on a 55,393-sample spliced test set while running at 3.65 ms per 16‑frame clip with a 150.4 MB checkpoint.
By Tamoghna Chakraborty, Md Nurul Absur, Sourya Saha, Saptarshi Debroy