AdaRoboVLG is a Vision‑Language‑Grasp framework that separates a generalizable base grasp policy from task‑specific understanding. The base policy generates and evaluates physically feasible grasp candidates using kinematic mapping and force‑closure stability, while foundation‑model modules supply composable spatial, cognitive, and temporal priors that adapt grasp synthesis to different robotic hands and environments without retraining. Experiments show strong cross‑hand generalization, effective handling of diverse grasping challenges, and functional grasping in cluttered, dynamic settings.
By Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang
The paper introduces the Temperature Scaling Attack (TSA), a training‑time method that degrades model confidence calibration while keeping predictive accuracy largely intact. TSA injects temperature scaling with a learning‑rate coupling during local federated training, shifting confidence scores and causing significant calibration errors (e.g., a 145% increase on CIFAR‑100) with less than a 2% drop in accuracy. The authors provide a convergence analysis for non‑IID settings and demonstrate TSA’s effectiveness across three benchmarks, robust aggregation, and post‑hoc calibration defenses, highlighting its impact on mission‑critical systems such as healthcare verification and autonomous driving.
By Kichang Lee, Jaeho Jin, JaeYeon Park, Songkuk Kim, JeongGil Ko
FailBench is a new benchmark for robot failure detection, containing 2,197 manipulation attempts from 14 public sources, with 75% of failures occurring naturally. The study evaluates 13 vision‑language model (VLM) detectors, finding the best model achieves only 0.77 mean balanced accuracy, and that fine‑tuned failure detectors often underperform general‑purpose VLMs. Performance varies with visual evidence, excelling when object motion is observable but dropping to near chance on contact‑intensive assembly tasks, and input‑level cropping of outcome‑relevant regions improves the top detector by 2.4 percentage points.
By Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan
The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, allowing attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate these methods, especially for sparse, real‑world differences, with open‑weight models to ensure reproducibility.
By Julian Truetsch, Felix Hauser, Christoph Stiller, Frank Bieder
The paper introduces a decentralized, vision-based system using multiple quadrotors equipped with a single RGB camera for monitoring wildlife. It emphasizes scalability, low bandwidth, and minimal sensor requirements, enabling robust identification and tracking of large species in natural habitats. The authors present novel coordination and tracking algorithms that operate without centralized communication, and validate the approach with real-world field experiments.
By Makram Chahine, William Yang, Alaa Maalouf, Justin Siriska, Ninad Jadhav, Daniel Vogt, Stephanie Gil, Robert Wood, Daniela Rus
LightEMMA is a longitudinal evaluation framework that tests the autonomous driving performance of vision‑language models (VLMs) without fine‑tuning or prompt engineering. Using this protocol, the authors evaluated 15 VLMs from five major families on the nuScenes prediction benchmark and found that larger, more capable models do not consistently outperform earlier generations. The study identifies common failure modes such as overreliance on historical actions and difficulty reconciling conflicting visual cues, underscoring the need for domain‑specific adaptation to enhance VLM safety in autonomous driving.
By Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
IRWOZ 2.0 is a refined dialogue dataset for industrial human‑robot interaction, expanding to 390 dialogues across four domains—Assembly, Delivery, Position, and Relocation. The dataset was improved using large language models (Mistral and Claude‑3.5) for generation and quality refinement, including manual corrections and automated typo removal. Benchmark tests show a substantial boost in dialogue state‑tracking performance, with GPT‑2’s BLEU‑4 score rising from 0.1651 to 0.5604 compared to the original IRWOZ.
By Chen Li, Dimitrios Chrysostomou
Skyfall-GS is a hybrid framework that generates large‑scale, city‑block‑sized 3D urban scenes by combining satellite imagery for coarse geometry with open‑domain diffusion models for detailed appearance. It uses a curriculum‑driven iterative refinement to improve geometric completeness and photorealistic textures, eliminating the need for costly 3D annotations. Experiments show that Skyfall‑GS achieves better cross‑view geometry consistency and more realistic textures than existing methods.
By Jie-Ying Lee, Yi-Ruei Liu, Shr-Ruei Tsai, Wei-Cheng Chang, Chung-Ho Wu, Jiewen Chan, Zhenjun Zhao, Chieh Hubert Lin, Yu-Lun Liu
RoboTok is an internet‑scale data engine that retrieves human manipulation videos from the web to train dexterous robot policies. It learns a latent motion space from 3D hand trajectories in actor‑centered reference frames, allowing manipulation behaviors to be compared across different viewpoints, scenes, and occlusions while remaining compact for efficient search. Experiments show RoboTok retrieves more relevant demonstrations and improves downstream robot task success compared to existing retrieval methods.
By Howard Qian, Yiting Chen, Yunfei Xie, Kejia Ren, Podshara Chanrungmaneekul, Gaotian Wang, Bowen Wen, Chen Wei, Kaiyu Hang
The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to improve goal‑conditioned robotic planning. By preventing latent collapse and grounding representations in physical configuration, the model achieves top success rates on tasks such as TwoRoom, PushT, and OGBench‑Cube, outperforming the baseline LeWorldModel. Ablation studies confirm that state alignment consistently boosts planning success over inverse dynamics alone across all four benchmark tasks.
By Muyuan Liu (GENISOM AI, Beijing, China), Yue Huang (GENISOM AI, Beijing, China), Zheng Liang (GENISOM AI, Beijing, China), Xiang Gao (GENISOM AI, Beijing, China)
The paper introduces Evidence‑Gated Regularization (EGR), a modality‑agnostic training objective that mitigates modality entanglement in Vision‑Language‑Action (VLA) policies. EGR uses per‑frame, per‑sensor task‑relevance signals to enforce invariance on low‑evidence sensors and single‑sensor sufficiency on high‑evidence ones, adding no inference‑time overhead. Evaluations on a BEHAVIOR‑1K benchmark and two real‑robot setups (bi‑manual Kinova arms with RGB cameras and a single‑arm MELFA ASSISTA with vision and GelSight tactile sensors) show significant improvements in success rates across various corruption and fallback scenarios.
By Yue Yang, Diego Romeres, Chiori Hori, Gedas Bertasius, Daniel Szafir, Siddarth Jain
The paper introduces ORMOT, a new task that extends Referring Multi‑Object Tracking to omnidirectional 360° imagery, ensuring full scene context for language‑guided tracking. It presents ORSet, a dataset of 27 omnidirectional scenes with 848 language descriptions and 3,401 annotated objects, and introduces ORTrack, an LVLM‑driven framework that performs zero‑shot detection and robust cross‑frame association. Experiments on ORSet show that ORTrack achieves state‑of‑the‑art performance, establishing a strong baseline for future research.
By Zihan Zhou, Sijia Chen, Yanqiu Yu, En Yu, Wenbing Tao
The paper introduces LaPla, a Vision‑Language‑Action framework that uses a latent‑aligned planning approach to convert discrete semantic reasoning into continuous, physics‑constrained driving actions. It employs a residual VQ‑VAE to encode vehicle kinematics into a structured latent space, then projects multimodal inputs—images, past actions, and text—directly into this latent space, allowing a frozen decoder to generate physically plausible trajectories without quantization errors. Experiments on nuScenes and NVIDIA AlpaSim show LaPla reduces long‑horizon L2 error by 15.52% and improves closed‑loop success rates by 33.34 percentage points while cutting inference latency.
By Ruoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu, Zewei Yang, Yipeng Zhu, Xiaolong Wang, Jun Ma
The paper reviews 194 papers on using artificial intelligence to optimize data center energy use, coding 63 of them. It finds that most control studies validate only in simulation, none consider water withdrawal or embodied carbon, and savings estimates overlap across methods, preventing ranking. The authors propose CLEAR‑DC, a framework that links control and workload demand through elasticity, reports net benefits, and records energy, carbon, water, embodied share, and validation venue.
By Mohammed Basharath Ullah, Summaiya Unnisa Begum, Mohammed Nadeem Ullah
CoMAP introduces a framework that jointly evolves textual world models and agent policies through a closed‑loop interaction. At each decision step the world model forecasts future state feedback for candidate actions, while the agent reflects on the reliability of this feedback to refine its action. The resulting on‑policy trajectories are used to self‑distill and update the world model, improving prediction accuracy and long‑horizon decision‑making across embodied planning, web navigation, and tool‑use benchmarks.
By Youwei Liu, Jian Wang, Hanlin Wang, Wenjie Li
AnyBox is a zero‑shot framework that estimates the full 9DoF pose (6D pose plus 3D dimensions) of boxes from a single RGB‑D image, leveraging the geometric regularity of boxes. It alternates between pose and scale estimation, using a binary search guided by the discrepancy between a reprojected template and the observed mask, and employs a depth‑consistency filter and an early‑stopping rule to prune implausible hypotheses. On public benchmarks and a warehouse dataset, AnyBox improves detection AP by up to 36 points and boosts robotic box‑shelving success by 28%.
By Yintao Ma, Sajjad Pakdamansavoji, Charles Eret, Rui Heng Yang, Xuan Zhao, Yingxue Zhang, Tongtong Cao, Amir Rasouli
The paper introduces GEO Defender, a two‑stage defense system designed to protect generative search engines from malicious Generative Engine Optimization (GEO) attacks that rewrite web documents to manipulate generated answers. GEO Defender comprises a Shield Reranker, which learns a defensive residual to demote GEO‑rewritten documents while maintaining relevance, and a Training‑Free Shield Generation component that creates a natural‑language library guiding the target LLM’s source usage during inference. Experiments on both closed‑source and open‑source large language models show that GEO Defender dramatically lowers attack success rates from 50.32% to 6.20%, preserves over 94% of benign evidence usage, and maintains answer quality while generalizing to unseen attacks.
By Haozhang Li, Yangguang Shao, Xinjie Lin, Zhong Guan, Mi Zhou, Junzheng Shi
The paper presents a complete characterization of when two deep ReLU networks realize the same function, showing that this occurs iff one can be transformed into the other using a set of axioms from many‑valued logic. It introduces a symbolic calculus that maps networks to substitution graphs, proves a completeness theorem linking equivalent formulas, and provides an algorithm to reconstruct networks from these graphs. The framework yields a new compositional normal form for MV logic that preserves the algebraic structure of deep ReLU networks.
By Yani Zhang, Helmut B\"olcskei
The article discusses how current world models, while achieving high predictive likelihood and visual fidelity, often fail to preserve the evidence needed for safe decision-making in embodied systems. It identifies three structural mismatches—likelihood versus risk, prediction versus intervention, and finite-horizon prediction versus accumulated consequences—and proposes the Risk‑Informed World Model (RIWM) as a decision‑centric framework. RIWM emphasizes consequences, intervention, epistemic uncertainty, and recoverability, integrating decision‑relevant representation, counterfactual reasoning, safety‑critical episodic memory, and runtime safety assurance to better support safety‑critical embodied systems.
By Kailang Ma, Heye Huang, Inhi Kim, Kitae Jang
The study examines how the speed and accuracy of an AI teammate—Fast/Less-Accurate (FLA-AI) versus Slow/Accurate (SA-AI)—affect performance in a collaborative Brain‑Computer Interface (cBCI) team during a virtual reality drone search task. Fast AI leads to instant, blind compliance and a sharp drop in human accuracy, while Slow AI induces delayed cognitive conflict that ultimately allows teams to recover and achieve perfect accuracy. A 2D Adaptive Riemannian Oracle and Hybrid Fusion techniques were used to adaptively capture and integrate these timing-dependent signals, improving team performance in both scenarios.
By Christopher Baker, Stephen Hinton, Akashdeep Nijjar, Riccardo Poli, Caterina Cinel, Tom Reed, Stephen Fairclough