The paper introduces SCARF, a real‑time velocity‑invariant representation for event cameras that preserves raw events while handling fast motion, stationary scenes, and independently moving objects. Unlike previous methods that convert events into image‑like forms and lose temporal detail, SCARF maintains temporal information and achieves state‑of‑the‑art performance in computational efficiency and representation quality.
By Mikihiro Ikura, Luna Gava, Jiahang Wu, Chiara Bartolozzi, Arren Glover
AnyviewMeter is a geometry‑conditioned adaptation framework that fine‑tunes pretrained robotic reward models to account for camera viewpoint and occlusion. It uses low‑rank fine‑tuning, token‑aligned Plücker rays, and synchronous block attention to fuse synchronized multi‑view observations. Experiments on PickCube and simulated manipulation tasks show significant reductions in mean absolute error and improved temporal ordering compared to single‑view or RGB‑only fine‑tuning.
By Yuang Tu, Runjia Tan, Yujie Yan, Jinghan Hu, Chen Lv
Scene-Q is a confidence‑aware, coarse‑to‑fine querying framework for indoor mobile robots that grounds natural‑language queries in a 3D map. It normalizes encoder scores with temperature scaling and only invokes a reasoning VLM for low‑confidence cases, while high‑confidence queries are answered by fast retrieval. The method improves open‑vocabulary 3D instance segmentation on ScanNet200 and natural‑language 3D instance retrieval on real‑world reconstructions, especially for spatial and relational queries, while maintaining a substantial fraction of queries on the fast path.
By Juno Kim, Yesol Park, Hye-Jung Yoon, Byoung-Tak Zhang
SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.
By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo
VABench is a benchmark that tests general‑purpose multimodal large language models (MLLMs) on embodied spatial intelligence by requiring them to observe, reason, act, and revise based on visual demonstrations and active perception. The benchmark includes 14 task families, a fixed model‑agnostic controller, and evaluates models on target localization, spatial relations, and long‑horizon composition tracks without providing privileged object poses or learned action heads. Results show that while the best model achieves perfect target localization, overall task success remains modest, and active camera control and geometric transfer significantly influence performance.
By Zhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu
The paper introduces Feel‑WM, an off‑road navigation world model that incorporates proprioceptive data to predict both visual scenes and the robot’s physical sensations such as slip, tilt, and shake. By learning a future proprioceptive state and failure risk from the robot’s own experience, the model can evaluate planned trajectories using a separable score that balances goal similarity with predicted failure risk. Experiments on real and simulated off‑road data show that Feel‑WM outperforms visual‑only models in both open‑loop planning and closed‑loop navigation for wheeled and legged robots, and it successfully guides a Husky robot around rough terrain on mountain trails where an end‑to‑end policy fails.
By E-In Son, Dong-Wook Kim, Ji-Hoon Hwang, Kangsun Lee, Jisung Bae, Jung-Taak Kim, Seung-Woo Seo
Navi-Agent is a zero‑shot Vision‑Language Navigation in Continuous Environments (VLN‑CE) agent that builds a coordinate‑free spatial state from visual observations and motion history. It represents this state as a navigation topology with nodes as visual places and edges as motion transitions, enabling observation‑based self‑localization, task progress verification, and recovery. Experiments on a zero‑shot VLN‑CE benchmark and real‑world robot platforms demonstrate that Navi‑Agent achieves state‑of‑the‑art performance among geometry‑constrained methods while remaining competitive with geometry‑based approaches.
By Wenyuan Xie, Mengyang Hong, Yongzhong Wang, Yanbiao Ji, Yijin Zhou, Shaokai Wu, Shalayiding Sirejiding, Huayi Zhou, Yi-Chao Chen, Ma Ling, Yue Ding, Hongtao Lu
INSPECT is a system that learns how a robot should choose its camera view during assembly inspection by observing a smart‑glasses assistant that answers part queries and guides the user. It uses techniques such as Presence‑Invariant TwinSwap for object evidence calibration, claim‑indexed supervision to separate evidence needs from camera changes, and object‑centered calibration to adapt view preferences to robot poses. In experiments on gearbox assemblies and angle‑grinder recordings, INSPECT outperforms other non‑oracle policies, improving view utility and decision accuracy.
By Di Wen, Kailun Yang, Wenhao Guo, Yitian Shi, Junwei Zheng, Yufan Chen, Ruiping Liu, Jiale Wei, Rania Rayyes, Kunyu Peng
SalsaAgent is a multimodal embodied language model that generates expressive, full‑body salsa follower motions in response to a human leader and music. The approach treats partner interaction as nonverbal token passing, extending a large language model’s vocabulary to include discrete motion, pairwise relation, and audio tokens. A two‑stage token‑to‑diffusion pipeline, combined with full‑body and pairwise‑relation tokenizers and alignment with automatically derived text descriptions of skeleton dynamics, yields improved motion quality, spatial coordination, and music‑partner synchrony compared to prior baselines.
By Payam Jome Yazdian, Zoe Stanley, Angelica Lim
The paper introduces G2G, a method that leverages known intra-group geometry to estimate the relative 6-DoF pose between two image groups, a key problem in cross-sequence relocalization and multi-camera rig odometry. G2G keeps a frozen foundation model and adds three lightweight trainable modules—a perceiver resampler, a cross-group bridge with merged self-attention, and a multi-frame pose head—totaling about 32 M parameters, less than 6% of the full model. Evaluated on four diverse datasets covering indoor/outdoor simulation, real-world cross-season capture, and zero-shot sim-to-real transfer, G2G achieves state‑of‑the‑art accuracy on both pose estimation tasks while only requiring supervision from relative poses.
By Yufei Wei, Shuhao Ye, Chenxiao Hu, Yiyuan Pan, Dongyu Feng, Rong Xiong, Yue Wang, Yanmei Jiao
The paper introduces PreDE, a policy‑calibrated framework that predicts how post‑training quantization will degrade task performance in world action models (WAMs) before deployment. By calibrating two thresholds on a small development set, PreDE can accept, reject, or defer new quantization configurations based on offline action deviations, achieving 75% coverage of decisions that match closed‑loop outcomes. Experiments on five WAMs and real‑robot trials show that PreDE accurately identifies high‑deviation configurations and enables significant speedups and memory reductions without compromising performance.
By Jiuyi Xu, Jinjia Guo, Meida Chen, Jing Du, Yangming Shi
FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.
By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury
The paper introduces Latent-Centroid Steering (LCS), a single-pass classifier-free guidance method for vision‑language autonomous driving models. LCS replaces instance‑level residuals with class‑level latent shifts, projecting conditional representations toward precomputed command‑specific centroids to enhance command adherence. Experiments on Bench2Drive and nuScenes show that LCS cuts inference latency by about 50% while improving driving performance.
By Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao
REACT is a fully spiking state‑space model that processes raw event‑camera data one event at a time, avoiding temporal accumulation and its associated delay. It employs a complex‑valued spiking neuron (C‑SiLIF) whose dynamics are driven by the inter‑event interval, enabling continuous‑time state updates at microsecond resolution. Evaluated on gesture recognition and time‑to‑collision estimation, REACT achieves low latency (4.6 ms) and high accuracy, supports anytime prediction, zero‑shot transfer, and INT8 quantization, dramatically reducing energy consumption.
By Geoffroy Keime, Nicolas Cuperlier, Benoit R. Cottereau
The paper proposes using action‑similarity supervision to improve cross‑embodiment transfer in latent action models (LAMs). By training the similarity between latent actions to match the similarity of ground‑truth robot action sequences—without predicting the actions themselves—the authors reduce sensitivity to background noise and embodiment differences. Experiments on RoboTwin 2.0 show that this approach more than doubles cross‑embodiment success compared to predicting ground‑truth actions, especially when similarities are computed on end‑effector motion and compared across robots.
By Maxime Alvarez, Renzo Caballero, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo
The paper introduces InterTrack, a behavior world model that enables humanoid robots to perform robust whole-body tracking while interacting with varied terrain and objects. Using a Transformer architecture, InterTrack predicts actions, states, and behavior distributions conditioned on the environment, and it is trained with an automated pipeline that reconstructs 3D support geometry from retargeted motions. The system achieves an 81.3% success rate on terrain interaction, a 99.3% fall-recovery rate, and outperforms leading baselines in both free-space tracking and cross-terrain scenarios.
By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Borui Zhang, Bingyao Yu, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
The paper introduces an open‑vocabulary 3D object detection pipeline that uses a promptable segmentation model (SAM3) to generate instance masks from six surround‑view cameras. These masks are converted into metric 3D boxes, achieving up to 0.413 mAP/0.555 NDS without any training when supervised box geometry is borrowed at inference. The approach also improves a supervised LiDAR‑only detector by 0.034 mAP through a camera‑witness rule, demonstrating that measurement precision, not 2D detection, limits performance.
By \"Omer Faruk Deniz, Mustafa Taha Ko\c{c}yi\u{g}it
ParticleSplat is a self‑supervised, object‑centric learning framework that extends the Deep Latent Particles (DLP) model into 3D by representing scenes as latent particles mapped to 3D Gaussian splats. It jointly encodes multiple camera views into a shared 3D latent space, enabling unsupervised learning of object masks and controllable 3D scene editing such as moving objects by manipulating latent particles. Experiments on simulated and real‑world datasets demonstrate that this 3D representation improves performance on downstream robotic manipulation tasks.
By Lyuxing He, Daniel Guo, Elizabeth Terveen, Deepak Pathak, David Held, Tal Daniel
The paper introduces Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing (A‑RAM), a framework that transforms user intent and part files into execution‑ready plans for robotic AM. It uses a large language model to interpret manufacturing goals, a deterministic Planning Agent to generate search workflows, and domain tools to evaluate slicing, placement, inverse kinematics, trajectory timing, Joint‑6 jerk, and extrusion. Experiments on a six‑axis robotic‑arm AM cell show that A‑RAM can reduce maximum Joint‑6 jerk by up to 53.5 % and mean absolute Joint‑6 jerk by 48.3 %, while also shortening motion‑plan completion times and extrusion paths.
By Jingzhan Ge, Ruimin Chen, Azadeh Haghighi, Jiong Tang, Farhad Imani
The article introduces a mathematical model of the Motivated Emotional Mind cognitive architecture for embodied intelligent systems. It formalizes a re‑entrant loop that integrates feedforward processing, lateral interactions, and feedback pathways, along with representational selection mechanisms that govern adaptive responses. The model binds exteroceptive and interoceptive signals, bodily-motivational context, and memory traces into associative structures called semblions, enabling motivated learning that incorporates need thresholds, goal generation, and regulatory constraints.
By Wies{\l}aw L. Galus, Janusz A. Starzyk