The paper introduces LEAP-CBF, a safety filter that uses Least‑Effort Adversarial Potentials to quantify how much disturbance effort is needed to cause failure in nonlinear dynamical systems. LEAP serves as a control barrier function for the undisturbed system and can be combined with a robust safety filter that tolerates disturbances with bounded cumulative effort. The authors develop a deep reinforcement learning method to construct LEAPs and validate their effectiveness through simulations of multi‑agent systems and hardware experiments on a quadruped and quadrotors.
By Oswin So, Eric Yu, Chuchu Fan
The paper evaluates infrastructure‑based radar for highway traffic monitoring using the newly introduced DRaT dataset, which pairs radar data with drone‑derived ground truth. It reports that radar achieves 78 % precision and 57 % recall for vehicle detection, tracks vehicles with an IDF1 score of 0.699, and estimates space‑mean speed with less than 4 % error while underestimating density and volume by about 23 %. The study also highlights deployment considerations and releases the dataset for reproducible research.
By Tianheng Zhu, Woei-chyi Chang, Alamss Riaz, Sogand Hasanzadeh, Yiheng Feng
AWM‑VLA introduces a unified framework that embeds aligned world modeling directly into a diffusion‑transformer vision‑language‑action policy. By adding learnable future tokens aligned with vision‑language embeddings of future observations, the policy can anticipate long‑term consequences while generating actions. The method extends this with an object‑centric alignment objective and a principled weighting scheme, achieving up to 21% higher success rates on RoboCasa and humanoid tabletop benchmarks and producing object‑centric rationales preferred by human raters in 83% of cases.
By An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian
BronchoTop is a real‑time, RGB‑only framework that localises a bronchoscope within the bronchial tree without requiring patient‑specific CT scans or external sensors. It uses four modules—lumen detection and tracking, lumen‑branch label association, probabilistic scope location estimation, and switch verification—to estimate the scope’s position relative to a generic airway model. Evaluation on phantom, simulated, and real data shows state‑of‑the‑art accuracy, improving existing approaches by over 20% on real bronchoscopy sequences and providing the first publicly available framework with code and data for further research.
By Clara Tomasini, Ana Cristina Murillo, Luis Riazuelo
The paper introduces a privacy‑preserving approach for semantic segmentation that fuses high‑resolution depth with ultra‑low‑resolution RGB images. A joint 2D framework uses depth to guide RGB reconstruction and RGB‑D segmentation, while an end‑to‑end 2D‑to‑3D pipeline consolidates 2D features for 3D segmentation. Experiments on ScanNet demonstrate superior 2D and 3D performance compared to other privacy‑preserving methods, strong zero‑shot transfer to SUN RGB‑D and SceneNN, and reduced recoverability of sensitive data, with real‑robot tests showing effective object‑goal navigation.
By Xuying Huang, Swithinraj Moses Daniel, Sicong Pan, Sebastian Houben, Maren Bennewitz
The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.
By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
The paper introduces a laser‑tracker‑assisted hand‑eye calibration technique for camera‑equipped mobile robots. It integrates 3D metrology from a laser tracker with 2D camera observations, extending a prior method by chaining multiple calibration targets to relax constraints on robot and camera configurations. This generalized formulation enables calibration of the camera pose within the tracker‑localized robot coordinate system for a variety of mobile robot systems.
By Jan A. Rudolph, \"Oyk\"u Kandemir, Markus Ulrich
Behaviora is a preliminary conceptual architecture that represents both external and internal behavior of robots and agents in an addressable form. It defines a Behavior Episode composed of components derived from behavior taxonomies, each assigned a persistent Internet of Behaviors (IoB) Address. The architecture also includes a Style Profile to describe how behavior is expressed, an Experience Profile to capture internal state influencing execution, and a Behavior Compiler to translate these representations into platform‑specific actions.
By Gote Nyman
The paper introduces the Systemic Risk Index, an open pipeline and dashboard that aggregates evidence from 19 public AI benchmarks into four systemic‑risk categories defined by the EU GPAI Code of Practice. It evaluates 18 models using harm‑preserving perturbations and simulated deployment contexts, offering users the ability to switch between average and worst‑case aggregation and to trace each risk rating back to its benchmark evidence. The study finds that worst‑case scores can be 14 to 37 points lower than average scores, and that LLM judges agree with human graders at a level comparable to human‑human agreement.
By Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin
The paper discusses how robotic embodiment—sensing, kinematics, dynamics, geometry, actuation, and control—varies across robots and over time, and argues that general embodied intelligence must learn across these differences. It critiques current methods that engineer correspondences for short‑term gains, proposing instead that learning should discover representations that enable transfer across a broader range of embodiments as experience accumulates. The authors advocate for embodiment diversity as a scaling axis, broad learned priors as a complementary ingredient, and evaluations that better characterize embodiment gaps and transfer performance, linking practical cross‑embodiment learning to the scientific pursuit of physical intelligence that adapts with its embodiments.
By Bo Ai, Henrik I. Christensen, Hao Su
The paper introduces Fed-ADR, a coordinated attack framework where a malicious orchestrator server directs heterogeneous adversarial clients to adapt their gradient updates in real time, thereby evading existing federated learning defenses and drastically reducing global model accuracy. It also presents a lightweight detection mechanism that estimates true client gradients from historical data to spot coordinated attacks, and an in-situ recovery method that restores model performance without restarting training. Experiments on MNIST, Fashion‑MNIST, and CIFAR‑10 show the attack can drop accuracy from over 90% to below 10%, while the defense can recover accuracy to above 90% within a few rounds at a computational cost at least 20× lower than retraining from scratch.
By Mohamed Shaaban, Ahmed Abdelnaby, Mohamed Elmahallawy
The study investigates how verbal attunement and real‑time behavioral mimicry affect users’ perceptions of an embodied AI counselor in virtual reality. Participants interacted with a system that varied in verbal attunement (attuned vs. neutral) and behavioral mimicry (present vs. absent). Results indicated that verbal attunement most reliably increased perceived empathy, while mimicry had a marginal effect on perceived humanness and showed exploratory positive associations with empathy, positivity, and humanness, especially among female participants.
By Nathalia Gomez, Haig Shamlian, Omar Khan, Tiffany D. Do
The paper introduces Safety to Competence (S2C), a two‑stage reinforcement learning framework that first learns a safety filter and then trains a competitive task policy while embedding the filter. By separating safety synthesis from task learning, S2C reduces training complexity and prevents the policy from being exploited by adversarial attacks. Experiments on simulated touchdown games show that S2C achieves higher win rates, better Elo ratings, and lower exploitability than eight safe‑RL baselines, and hardware tests confirm its competence against a human opponent.
By Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu
InfiNoVA is a data‑augmentation framework that transforms synchronized multi‑camera demonstrations into a dense, geometrically consistent set of training views by reconstructing each manipulation trajectory as a time‑varying 3D Gaussian. The method renders novel observations from sampled camera poses while preserving the original state‑action pairs, improving frame‑level fidelity and temporal consistency compared to generative synthesis. Across four real‑world manipulation tasks, policies trained with InfiNoVA achieve 5.4× higher average success under unseen randomized viewpoints than VISTA‑based augmentation and 1.7× higher success than training on all five physical camera views.
By Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave
MemBodied introduces a fixed‑size episodic memory for Vision‑Language‑Action models, comprising an associative state that tracks interactions across policy calls and an episode anchor that stores a compact representation of the initial scene. By conditioning action generation on these memory components instead of raw past observations, MemBodied reduces context bloat and inference latency. In five memory‑dependent RMBench tasks, it outperforms stateless and vanilla recurrent policies by significant margins, and achieves a 90.6% success rate on the LIBERO‑Long suite, improving over the baseline by 5.4%.
By Tej Deep Pala, Navonil Majumder, Bryce Goh, Raphael Yee, Jianfei Yang, Liming Chen, Soujanya Poria
The paper introduces a language‑guided approach for robots to join human groups by predicting socially compliant joining poses. It uses recursive spectral partitioning to generate candidate group subsets, ranks them with a language‑conditioned image–geometry model, and then applies a goal predictor that incorporates human‑formation priors to produce a multimodal energy–orientation map of feasible robot poses. Experiments on various group scenarios show competitive grounding accuracy, superior joining‑pose prediction, and successful real‑robot demonstrations in both static and dynamic settings.
By Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu
Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.
By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
The paper investigates whether using non‑Gaussian priors improves fine‑tuning of large behavior models (LBMs) for robot imitation learning. Across more than 100,000 simulation rollouts and 1,250 hardware trials on diverse tasks, the authors find that non‑Gaussian priors do not yield better fine‑tuning performance than standard Gaussian priors, except possibly at very low data fractions. Diagnostic analyses reveal that encoder training dominates fine‑tuning outcomes, while prior choice has minimal impact.
By Chen Xu, Rishi Shah, Hadas Kress-Gazit, Haruki Nishimura, Masha Itkina
Teach-to-Crash is a closed‑loop testing framework that uses a dual‑LLM architecture to generate collision‑inducing scenarios for autonomous driving systems. A high‑reasoning Teacher LLM controls the search when collision metrics stagnate, while a low‑reasoning Student LLM produces simulator‑executable scenarios in JSON. In a CARLA case study, Teach‑to‑Crash achieved the highest collision hit rate (90.79 %), the shortest mean time‑to‑collision (18.31 s), and superior diversity and avoidability metrics compared to other methods.
By Zaid Ghazal, Khouloud Gaaloul, Bruce Maxim
The paper introduces TSGPD-IR, a network that disentangles weather-induced artifacts from true thermal signals in infrared images. It uses weather semantics and regional degradation severity to generate adaptive prompts, estimate severity without manual labels, and select appropriate expert modules for restoration. This approach aims to reduce artifacts and preserve weak thermal details across varying weather conditions.
By Xinyao Wang, Lijun He, Zhihan Ren, Fan Li