FuncBridge is a two‑stage framework that addresses functional generalization in robotics by decoupling functional reasoning from action execution. It learns to predict generalizable 2D keypoint trajectories from action‑free data and then grounds these trajectories into robot actions with limited demonstrations. Across a benchmark of ten tools and three functions—hitting, sweeping, and hooking—FuncBridge outperforms state‑of‑the‑art methods on unseen tools in both simulation and real‑world tests.
By Chuhao Zhou, Liquan Wang, Shuxin Cao, Xiangyu Chen, Yuxuan Hu, Boyu Ma, Animesh Garg, Jianfei Yang
Diffusion-2BC is a hybrid training method that combines a diffusion denoising objective with an auxiliary deterministic behavior‑cloning loss over a shared visual encoder. The auxiliary loss is used only during training, while inference remains diffusion‑based. Experiments on the Claw environment and CARLA navigation show that Diffusion‑2BC reduces mean mask‑distance error by about 10% compared to a diffusion baseline and by 85% compared to standard deterministic behavior cloning, and it enables the agent to travel farther and exhibit multimodal route choices.
By Bruno Maciel Machado, Eric Aislan Antonelo
The paper investigates whether memory-dependent control in simulated air‑hockey defense requires nonlinear recurrent dynamics. Using a DreamerV3 teacher, the authors show that a memoryless policy fails under temporary puck‑tracking loss, while a teacher with a reset recurrent state performs poorly, indicating the necessity of memory. They distil the teacher into compact recurrent policies with a 64‑dimensional state, employing a diagonal linear recurrence and optional low‑rank nonlinear innovation; the purely linear recurrent model matches the GRU baseline and teacher across all tested conditions, with no benefit from added nonlinear innovation.
By F. Olivia Fan, Oliver Obst
The paper presents a theoretical framework for human torque estimation-based task‑agnostic control of robotic exoskeletons, defining matched assistance as scenarios where the robot positively contributes to human movement. It designs the robot’s desired interaction torque to guarantee a lower bound on matched assistance probability across the entire torque distribution, including unseen data. Experimental validation on the ABLE upper‑limb exoskeleton shows that the strategy achieves smooth movement and reduced human effort across multiple tasks.
By Duy Hoang, Bastien Berret, Olivier Bruneau, Laurent Fribourg
The article outlines a Ph.D. research agenda aimed at creating provably efficient and practical algorithms for data‑driven sequential decision‑making under uncertainty. It focuses on reinforcement learning and multi‑armed bandits, targeting applications such as recommendation systems, computer networks, video analytics, and large language models. The work seeks to overcome limitations of existing methods—such as reliance on idealized models, lack of robustness to adversarial perturbations, and poor instance‑dependent performance—by developing algorithms that are more efficient, robust, instance‑adaptive, and generalizable to new environments.
By Zhiyong Wang
Fast‑TD‑MPC is a lightweight framework that blends fast policy execution with selective test‑time planning for data‑driven model predictive control. By routing between intuitive, low‑cost actions and deliberative trajectory optimization only when needed, it achieves competitive performance on 103 continuous control tasks while speeding inference up to four times. The method also retains robustness to external disturbances by falling back to planning when required.
By Yi Xian Goh, Sze Jue Yang, Hao Luan
FutureWorlds is a framework that learns robotic world models by generating alternative future predictions and using them as learning signals. It combines candidate construction, history maintenance, and learning from relative quality through a multimodal discrete autoregressive model and diverse beam search. The MemSPO algorithm further optimizes the world model by converting video trajectory rewards into group-relative advantages, leading to significant improvements in LPIPS scores across RT-1, BridgeV2, and RoboCasa datasets.
By Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang
The paper presents a framework for creating a scalable synthetic dataset of controllable video interactions by generating explicit start and end state images using image editing models. It introduces State‑Guided Sampling (SGS) to produce seamless videos anchored on these states, reducing artifacts seen in naive conditional generation. An automated evaluation system aligned with human judgments is also developed, and experiments demonstrate that fine‑tuning a base model on this dataset markedly improves its ability to generate plausible interactions.
By Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim
The paper introduces HierGF, a method for sparse‑view 3D reconstruction that uses a hierarchical geometry‑perception backbone to generate structured pseudo‑supervision from limited views. It incorporates a learnable confidence network to focus gradients on cross‑view consistent content and a geometrically consistent densification module to improve alignment and fill under‑sampled regions. This approach aims to overcome the challenges of few images and incomplete coverage that typically lead to overfitting and structural holes in existing methods.
By Bi'an Du, Zhimin Zhang, Daizong Liu, Baoquan Chen, Wei Hu
Ego2Act is a new benchmark that tests video generation models on goal‑directed, egocentric manipulation tasks. It contains 2,640 videos from 110 real‑world tasks, each requiring multiple steps of object manipulation to achieve a high‑level goal. The benchmark includes Ego2ActJudge, a reference‑free evaluation pipeline that better aligns with human judgments of task completion and physics plausibility.
By Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, Alham Fikri Aji
PhysicsLENS is a new dataset and benchmark designed to evaluate how well video generation models capture physical properties relevant to robotics. It consists of matched scenario pairs that share the same conditioning frame and task but differ in underlying physics, and it covers seven physical domains such as collision, gravity, and friction. The benchmark includes over 400 human-annotated labels from four video generation models, revealing that many plausible-looking videos still ignore the specified physical property.
By Isaiah Milkey, Som Sagar, Aditya Taparia, Xinyuan Liu, Jiqing Wen, Ransalu Senanayake
PAGER is a label‑free adaptation method that aligns partial, viewpoint‑dependent 3D observations with a frozen global semantic space. It uses matched‑point feature alignment and relational supervision to anchor partial features to their global counterparts while preserving similarity structure, all without altering the pretrained encoder or global probe. Experiments show PAGER outperforms label‑supervised PEFT on Sonata and Concerto, and achieves superior zero‑shot transfer from ScanNet to ScanNet++ compared to fully fine‑tuned Sonata.
By Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni
GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.
By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
LiteReality-Agent is an agentic system that reconstructs real indoor environments into realistic, articulated, and simulation-ready 3D scenes from RGB‑D scans. It treats reconstruction as a coding problem, where a coding agent iteratively edits a Python script (Room.py) using specialized tools, while an observe‑edit‑verify harness ensures evidence gathering, measurement, verification, layout optimisation, simulation readiness, and quality control. The system outperforms recent models like Astra and Fable in geometric accuracy, visual realism, and simulation compatibility, and its modular framework is positioned as a robust foundation for future agent‑driven reconstruction tasks.
By Zhening Huang, Yueyan Li, Johnathan Chiu, Xiaoyang Lyu, Matt Zhou, Yuxin Yao, Joan Lasenby, Shangzhe Wu
MapLightning introduces a new approach to online vectorized HD map construction that replaces dense bird’s‑eye‑view grids with a compact set of 1D learnable map tokens. Using a transformer‑based mapper that applies full self‑attention to combined image and map tokens, the method achieves efficient inference, lower memory usage, and improved global context without relying on camera projection parameters. Experiments on nuScenes and Argoverse 2 show state‑of‑the‑art accuracy, with a lightweight variant outperforming MapTRv2 by over 10 mAP and running 1.73× faster at 40+ FPS while using 53 % less memory.
By Shen Zheng, Anurag Ghosh, Mani Ramanagopal, Srinivasa Narasimhan
The paper introduces regularizers that enforce signal‑noise factorization (SNF) and signal‑signal factorization (SSF) during training of deep neural networks. Experiments on CIFAR‑100 show that SNF regularization improves classification accuracy, while SSF does not. On the BloodMNIST dataset with varying corruption levels, SNF yields even larger gains, and analysis reveals that SNF isolates noise into distinct subspaces, enabling projection of corruption‑induced directions and further accuracy improvements.
By Sakin Kirti, Joel Zylberberg
FocusGraph is a framework for selecting keyframes in egocentric long‑video question answering. It uses a lightweight Scene‑Graph LLM Selector to identify query‑relevant clips from graph‑based captions, then extracts keyframes with Patch‑wise Sparse‑Flow Retention (PSFR) before feeding them to a multimodal large language model for answer generation. The method achieves state‑of‑the‑art performance on FindingDory and HourVideo while reducing question‑time inference cost.
By Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin
SlotVLA proposes a slot-attention framework that models both objects and their relations for robotic manipulation, using a relation‑centric decoder and an LLM‑driven action translator. The authors introduce LIBERO+, a fine‑grained benchmark with object‑centric annotations, box and mask labels, and instance‑level temporal tracking to evaluate object‑relation reasoning. Experiments on LIBERO+ show that object‑centric and object‑relation slot representations reduce visual token counts while maintaining competitive generalization, offering a compact and interpretable foundation for multitask visuomotor control.
By Taisei Hanyu, Nhat Chung, Huy Le, Toan Nguyen, Yuki Ikebe, Anthony Gunderman, Duy Nguyen Ho Minh, Khoa Vo, Tung Kieu, Kashu Yamazaki, Chase Rainwater, Anh Nguyen, Ngan Le
DynamicVLA is a latency‑aware Vision‑Language‑Action model designed for dynamic object manipulation, featuring a compact 0.4B architecture and a convolutional vision encoder for efficient multimodal inference. It employs a continuous inference schedule that overlaps reasoning and execution, and a Latent‑aware Action Streaming mechanism that discards stale action prefixes to maintain action‑time alignment. The authors also introduce the Dynamic Object Manipulation (DOM) benchmark, comprising 200K synthetic episodes and 2K real‑world episodes, and demonstrate that DynamicVLA improves dynamic manipulation success in simulation and on real robots.
By Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu
The paper reviews one‑pixel attacks (OPAs), highlighting their extreme adversarial fragility across domains such as medical diagnosis, autonomous driving, biometrics, and quantum communication. It presents a PRISMA‑guided synthesis of studies from 2017 to 2026, offering a unified taxonomy that covers algorithmic foundations, black‑box evolutionary optimization, hybrid and program‑synthesis attacks, defence mechanisms, interpretability tools, and domain‑specific vulnerabilities. The review identifies dominant Differential Evolution strategies, emerging efficiency‑optimized and saliency‑guided methods, and gaps in dataset diversity, transferability, and evaluation standards, while proposing future research directions and a regulatory framework for robustness testing and AI security governance.
By Mirza Niaz Morshed, Md. Masudul Islam, Galib Muhammad Shahriar Himel, Md. Aslam Uddin, Hui Liu, Md. Shafiqul Islam