Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,858 stories · RSS feed

arXiv Machine Learning
Sep 22

Connectivity-Aware Exploration of Robotic Grasp Spaces

The paper investigates the multiscale connectivity structure of successful robotic grasps in SE(3) and finds that these sets exhibit heterogeneous yet reproducible connectivity across objects. It proposes a connectivity‑aware sampling strategy that prioritizes bridges, frontiers, boundary extensions, and geometric novelty, which recovers grasp connectivity more efficiently than random or farthest‑point sampling. Experiments also show that connectivity information can improve subsequent grasp discovery and transfer to unseen objects, indicating that the spatial organization of viable actions offers valuable guidance for exploration.

By Maksim A Kazanskii
arXiv Machine Learning
Sep 22

Beyond Appearance Shifts: Task-Semantic Action Calibration for VLA Models

The paper introduces BAS‑VLA, a task‑semantic action calibration framework for vision‑language‑action models that addresses two failure modes: unnecessary action drift under appearance changes and insufficient behavioral change under semantic alterations. BAS‑VLA uses a breaking‑centered calibration core and a selective evidence‑gated preserving auxiliary to maintain performance on clean and semantics‑preserving conditions while suppressing stale‑task behavior. Experiments on OpenPI‑pi0.5 and LIBERO‑Object Milk‑Swap show high success rates on clean and preserved tasks, a dramatic drop under target‑object swaps, and improved robustness to style shifts from 42% to 70% without harming clean performance.

By Shuaijun Liu, Feiyang You, Chengyu Wu, Shuyang Hao, Chenglong Zhang, Jingyao Cai, Xingwei Chen, Ningxin Su
arXiv Computer Vision
Sep 22

G6D: Geometric Learning-Free RGB-D 6D Pose Solver for Robotic Manipulation

G6D is a learning‑free, geometry‑driven RGB‑D 6D pose solver designed for robotic manipulation. It generates pose hypotheses via template‑based geometric matching and refines them using silhouette and depth consistency, requiring only an RGB‑D observation, an object mask, camera intrinsics, and a CAD model. The method offers adjustable accuracy‑computation trade‑offs, can run on CPU without GPUs, and has shown strong performance on LineMOD and BOP19 datasets, as well as in real‑world pick‑and‑place experiments.

By Yixuan Liang (Tsinghua University), William Chen (Sapient Intelligence), Yunan Wang (Tsinghua University), Jizhou Yan (Tsinghua University), Zhao Jin (Tsinghua University), Changling Liu (Sapient Intelligence), Chuxiong Hu (Tsinghua University)
arXiv Computer Vision
Sep 22

Active Spatial Inspection for Effective and Efficient Embodied Exploration

The paper introduces the ACE framework, which uses spatially explicit inspection to guide embodied exploration. By combining evidence‑grounded perception with exposure‑informed movement, ACE provides a spatially resolved decision paradigm that improves cue assessment and movement direction. Experiments show ACE boosts navigation task success by 18.0% and exploration efficiency by 10.3% over previous baselines.

By Wenbin Wang, Xiang Bai, Yizhao Wang, Hang Sun, Dong Ren, Jie Qin, Qingquan Li, Bing Wang
arXiv Computer Vision
Sep 22

Cognitive Action Reasoning for Proactive Robots from Human-Centered Multimodal Observations

The paper introduces ProAction, a multimodal dataset of 10,000 samples comprising visual, audio, and text inputs across 12 daily-life scenarios, designed to support the Proactive Robot Action Reasoning (ProRobo) problem. It presents a two-stage human-in-the-loop annotation pipeline that incorporates appraisal and Theory-of-Mind considerations to generate cognitively grounded high-level action labels. The authors benchmark multimodal large language models and propose MMC2Act, showing that training on ProAction significantly improves proactive action reasoning compared to general-purpose models.

By Zhihao Gu, Kechao Zhu, Yuanfeng Wu, Mohan Liu, Ankit Kumar Shaw, ChenDong Hong, Xuanyu Chen, Dengchen Mei, Xu Tianyi, Lin Wang
arXiv Machine Learning
Sep 22

CSC: Calibrated Simplicity for Conflict-Aware Social Bot Detection in the LLM Era

The paper introduces CSC, a calibrated-simplicity framework for detecting social bots in the era of large language models. CSC combines a simplified prototype-guided graph expert, calibrated simplex-constrained fusion, and a lightweight inconsistency expert to address modality conflict between semantic and structural signals. Experiments on TwiBot-22, TwiBot-20, and MGStBot-large demonstrate that CSC improves calibrated decision quality while maintaining competitive performance across benchmarks.

By Yipeng Qian, Pengjie Zhao, Chaoxi Niu
arXiv Machine Learning
Sep 22

D-JEPA: A Decision-Aligned Latent World Model

arXiv:2609.24749v1 Announce Type: cross Abstract: Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate wil...

By Shuaijun Liu, Chengyu Wu, Qifu Wen, Feiyang You, Chenglong Zhang, Shuyang Hao, Xi Lin, Ningxin Su
arXiv Computation and Language
Sep 22

The Role of AI in Online Reviews

arXiv:2609.22198v1 Announce Type: new Abstract: The rapid adoption of large language models (LLMs) creates new opportunities for strategic content generation on online platforms, including potentiall...

By Valeria Lerman, Oren Rigbi, Yaniv Dover