DeliveryGym is a 3D reinforcement learning environment that simulates continuous courier shifts, integrating multimodal tool interaction, persistent world dynamics, and trajectory‑based rewards derived from simulator events. It allows agents to learn how their decisions affect time, energy, and money across an entire shift, and it adapts future training shifts to the policy’s weaknesses while keeping evaluation fixed. Experiments on six models and 13 city maps show a significant gap between task execution and optimal sequencing, with RL improving Qwen3‑VL‑4B’s net income by 54.3% and adaptive training boosting test income by 16.5% over uniform sampling.
By Haoqiang Kang, Yiming Zhang, Yiyang Guo, Chuying Li, Jianzhi Shen, Tianruo Rose Xu, Xiaokang Ye, Lianhui Qin
The paper investigates whether closed‑loop robot software generated and refined by a coding agent can be reused to acquire policies for new tasks. For each source task, the agent creates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applied to new tasks, the agent uses these archived implementations, additional demonstrations, and execution feedback to produce a final policy that runs without further model calls, achieving higher success rates than starting from scratch or from unoptimized source code.
By So Kuroki, Yujin Tang
MaskHarness-WAM is an instance‑grounded harness designed to extend local visuomotor control to long‑horizon robot manipulation tasks. It links high‑level task planning with low‑level policies by generating and verifying target masks at each subtask boundary, enabling continuous target tracking and task progress assessment. Experiments on a real robot show that this approach outperforms limited‑horizon policies for sequential multi‑object manipulation.
By Zitai Huang, Taiyi Su, Jian Zhu, Jianjun Zhang, Chong Ma, Tianbin Liu, Weiyi Lu, Yi Xu, Hanli Wang
The paper introduces VLN on the Fly, an onboard vision‑language navigation stack for aerial robots that separates grounding, planning, and control into inspectable stages. A quantized vision‑language model grounds instructions to a coarse image cell, depth estimation lifts this to a 3D goal, a fast B‑spline planner generates a feasible trajectory, and a pretrained reinforcement learning policy translates the trajectory into motor commands. In controlled indoor flights, the stack achieved the target in 13 of 15 trials with a mean goal error of 5.72 cm and 39.3% GPU utilization, and successfully tracked collision‑free trajectories in cluttered environments.
By Marco S. Tayar, Felipe Tommaselli, Gianluca Capezutto, Pedro Antonio Rabelo Saraiva, Pedro H. V. de Freitas, Lucas Kido, Guilherme Sonego, Ricardo V. Godoy, Marcelo Becker
The paper introduces an automated pipeline that converts non‑critical driving scenes into safety‑critical scenarios by integrating computer vision, Large Language Models (LLMs), and Augmented Reality (AR). It detects and tracks road users, extracts safety features such as distance, velocity, motion direction, and Time‑to‑Collision (TTC), and evaluates scene criticality. Safe scenes are then modified by an LLM, which generates realistic collision‑inducing objects and behaviors that are overlaid onto the original scene using AR, achieving 97.52% safety classification accuracy on the nuScenes dataset and producing realistic scenarios like pedestrian crossings, rear overtaking vehicles, and sudden‑stop events.
By Noura Fady, Farah Khaled, Catherine M. Elias
TouchSight is a monocular egocentric vision framework that predicts dense full-hand contact forces from video. It uses 500 hours of pressure‑glove recordings and hand‑object interaction data, and introduces TwinTouch‑20H, a dataset of 20 hours of paired visual data where generative models render gloved recordings as bare‑hand observations while preserving tactile labels. The system outperforms prior methods on OakInk2, generalizes qualitatively to natural bare‑hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales.
By Danyan Zhou, Jinxuan Lu, Jiawei Lin, Tianxing Chen, Chuqiao Lyu, Wenbo Ding
The paper introduces Sampling-Guided Policy Search (SGPS), a method that combines sampling-based model‑predictive control with first‑order policy gradients to accelerate visual policy learning for locomotion and manipulation tasks. SGPS starts with behavior cloning from sampled actions and then alternates between sampling‑based refinement and short‑horizon policy updates under varied initial states and dynamics. The approach is demonstrated on simulated Unitree Go2 and G1 robots, learning tasks such as obstacle traversal and bimanual carrying, and the distilled policies transfer zero‑shot to a real Go2 robot using onboard depth perception.
By Yilang Liu, Haoxiang You, Qian Wang, Daniel Rakita, Ian Abraham
HIL-UMI is a policy-guided Universal Manipulation Interface that enables robot‑free, human‑in‑the‑loop post‑training of vision‑language‑action models. By querying the current policy during handheld demonstrations and using an Energy Score to detect out‑of‑distribution states, it selectively collects new data and refines a progress‑based advantage estimator. The updated estimator then drives advantage‑conditioned behavioral cloning, improving performance on long‑horizon and precise manipulation tasks while reducing per‑frame collection time compared to HG‑DAgger.
By Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao, Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang, Ruodai Li, Hui Shen, Hao Dong
GeoAAC introduces a geometry-based adaptive action chunking technique for Vision‑Language‑Action policies, dynamically adjusting the action horizon based on the reliability of current action predictions. By leveraging the geometric variation in Flow Matching denoising trajectories, GeoAAC constructs a horizon‑wise geometric profile that determines the action horizon during a single generation without extra training. Experiments on LIBERO, LIBERO‑Pro, RoboCasa365, and real‑world manipulation tasks demonstrate consistent gains over fixed‑horizon baselines, achieving up to 8.7 percentage points improvement in simulation and raising real‑world success rates from 53.3% to 74.4%.
By Xin Chen, Sen Chen, Yujuan Ding, Jian Liu, Guoqing Wang, Wei Ye, Heng Tao Shen, Yi Bin
The paper introduces SafeHarness, an obstacle‑aware framework that improves the safety of coding agents for robot manipulation. By decomposing tasks into route planning and contact execution, the harness enables the agent to prioritize collision avoidance, achieving 71.9% task success and 87.5% collision avoidance—significantly better than prior methods. The study demonstrates that safety constraints can be effectively integrated into language‑model‑driven robot controllers.
By Bingxin Xu, Yuzhang Shang, Zhen Dong, Emilio Ferrara
AntiGrounding is a visual action-selection framework that turns short robot trajectories into both executable motion plans and rendered prompts for vision‑language model evaluation. After filtering for feasibility, each trajectory is scored on safety, task alignment, efficiency, and physical plausibility using structured multi‑view visual question answering, and the best trajectories are refined and validated by a digital twin before real‑world execution. In eight real‑world manipulation tasks, the system achieved a 71.25% success rate with a single GPT‑6 Astra evaluator, outperforming baseline methods.
By Wenbo Li, Yiteng Chen, Wenhao Li, Qingyao Wu
Sim-and-Human Co-training (SimHum) is a method that combines simulation and human demonstration data to train bimanual manipulation policies. It first extracts kinematic priors from simulation and visual priors from human observations, then fine‑tunes on a small real‑robot dataset. With only 80 real‑robot episodes per task, SimHum achieves 62.5% success on out‑of‑distribution scenes across four tabletop tasks, outperforming real‑only training by 53.7% and improving the best single‑source baseline by 35.0% in a matched‑time study.
By Kaipeng Fang, Weiqing Liang, Yuyang Li, Ji Zhang, Pengpeng Zeng, Heng Tao Shen, Jingkuan Song, Lianli Gao
StarVLA-α is a streamlined Vision‑Language‑Action (VLA) model that reduces architectural and pipeline complexity to facilitate systematic study of VLA design choices. By employing a strong VLM backbone and minimal design, it achieves competitive performance across multiple benchmarks (LIBERO, SimplerEnv, RoboTwin, RoboCasa) and outperforms the baseline π₀.₅ by 20% on the RoboChallenge benchmark. The authors plan to release the code to support future VLA research.
By Jinhui Ye, Ning Gao, Senqiao Yang, Jinliang Zheng, Zixuan Wang, Yuxin Chen, Pengguang Chen, Yilun Chen, Shu Liu, Jiaya Jia
This study investigates whether large language models (LLMs) can reliably answer scientific questions and how susceptible they are to manipulation by fringe scientific material. The authors modified custom LLMs to prioritize knowledge from selected fringe papers on the Fine Structure Constant and Gravitational Waves, then compared their responses with those of domain experts and standard LLMs. The altered models produced fluent, convincing answers that contradicted scientific consensus and were difficult for non-experts to detect as misleading, demonstrating that LLMs are vulnerable to manipulation and cannot replace expert judgment.
By Harry Collins, Hartmut Grote, Paul Newbury, Patrick Sutton, Simon Thorne
EmbodiedMind introduces a three-stage training paradigm for embodied foundation models that tackles inefficient sample use, task imbalance, and credit assignment in long-horizon planning. The stages—Rejection Sampling-based Fine‑Tuning, Iterative Rejection GRPO, and Trie‑GRPO—filter low‑informative data, balance task difficulty, and use action prefix trees for step‑level advantage estimation. This approach yields a state‑of‑the‑art average performance of 70.02% across 18 benchmarks, notably improving long‑horizon task planning accuracy.
By Feifan Wang, Zongbing Zhang, Yu Zhang, Lingfeng Wang, Yurui Zhu, Jin Deng, Mingliang Zhang, Zhengguang Gao, Yongcheng Wang, Jin Xu, Ri Yang
Delphi Scanner is a static malware detection system for Windows PE files that balances efficiency and interpretability. It employs a convolutional neural network to model Windows API sequences and a rule‑based interpretation layer to map APIs to high‑level malicious capabilities. Tested on over 190,000 PE files, it achieves 95.35% accuracy with a 1.53 MB model, and demonstrates robustness against out‑of‑distribution samples and adversarial manipulations.
By Bijied Brahimi, Vincent Cohadon, Gabriel Glazman, Rayan Al Mohaize, Omran Berjawi, Rida Khatoun
MILER is an end‑to‑end reinforcement learning framework that achieves zero‑shot sim‑to‑real transfer for autonomous driving in unstructured environments. It uses a custom semantic mid‑level representation (MLR) simulator for offline training, and during deployment it processes real camera and LiDAR data with BEVFusion to produce a compatible bird’s‑eye‑view representation. The policy’s actions are applied via a trajectory‑alignment strategy, allowing the system to drive 17.3 km on a 3.0 km test track without human intervention, all running on a Jetson AGX Orin.
By Thomas Steinecker, Denis Trescher, Alexander Bienemann, Thorsten Luettel, Mirko Maehlisch
OPTED is a method for on‑policy fine‑tuning of end‑to‑end driving models that separates reinforcement learning from the policy update. A privileged teacher trained with RL on vectorized inputs (HD‑maps and bounding boxes) supervises the pre‑trained student during closed‑loop post‑training. Applied to the camera‑based models TransFuser and VaVAM in AlpaSim, OPTED boosts driving scores by 1.6× and 9.5×, respectively, while requiring roughly three orders of magnitude fewer simulator interactions than direct RL post‑training.
By Damiano Da Col, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, Christos Sakaridis
The paper introduces WZPlanner, a new dataset and model for safe autonomous driving in work zones. The dataset, WorkZonePlan, contains over 149,000 synthetic and 5,000 real-world samples with 3D annotations for lane and work zone boundaries, plus 228 evaluation routes in CARLA. The proposed BoundaryFormer (BF) and its enhanced BF++ variants jointly predict lane/work‑zone boundaries and driving trajectories, achieving higher Driving Scores while being significantly smaller than competing models.
By Nishad Sahu (Raj), Changzhong Qian (Raj), Guangzhou Cai (Raj), Shounak Sural (Raj), Ragunathan (Raj), Rajkumar
The paper surveys the evolution of smart glasses into AI smart glasses, framing them as wearable intelligence platforms that integrate egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world constraints for personalized assistance. It organizes the discussion into four dimensions: hardware foundations, wearable intelligence, interaction design, and application scenarios across healthcare, accessibility, learning, daily life, tourism, and industry. The authors identify five cross-cutting research challenges—next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models—to guide future work.
By Xu Yuan, Yi Wang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Shanru Lin, Guoliang Xing, Hongxia Yang, Jiannong Cao, Qing Li, Wenqi Fan