arXiv:2605.11151v3 Announce Type: replace
Abstract: Offline-to-online reinforcement learning (RL) improves sample efficiency by leveraging pre-collected datasets prior to online interaction. A key ch...
By Andrew Choi, Wei Xu
Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument...
3D Gaussian Splatting (3DGS) can reconstruct a captured scene photorealistically, but the resulting representation does not by itself support physical interaction. Robot simulation instead requires ob...
Robots operating around pedestrians often reason over a finite set of predicted human futures. Repeated online updates can concentrate this limited prediction budget on dominant destinations and leave...
Marketers now deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of human-panel cost. However, this procedure assumes t...
Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects in 2D and reconstruct them independently struggle i...
The paper investigates the multiscale connectivity structure of successful robotic grasps in SE(3) and finds that these sets exhibit heterogeneous yet reproducible connectivity across objects. It proposes a connectivity‑aware sampling strategy that prioritizes bridges, frontiers, boundary extensions, and geometric novelty, which recovers grasp connectivity more efficiently than random or farthest‑point sampling. Experiments also show that connectivity information can improve subsequent grasp discovery and transfer to unseen objects, indicating that the spatial organization of viable actions offers valuable guidance for exploration.
By Maksim A Kazanskii
The paper introduces BAS‑VLA, a task‑semantic action calibration framework for vision‑language‑action models that addresses two failure modes: unnecessary action drift under appearance changes and insufficient behavioral change under semantic alterations. BAS‑VLA uses a breaking‑centered calibration core and a selective evidence‑gated preserving auxiliary to maintain performance on clean and semantics‑preserving conditions while suppressing stale‑task behavior. Experiments on OpenPI‑pi0.5 and LIBERO‑Object Milk‑Swap show high success rates on clean and preserved tasks, a dramatic drop under target‑object swaps, and improved robustness to style shifts from 42% to 70% without harming clean performance.
By Shuaijun Liu, Feiyang You, Chengyu Wu, Shuyang Hao, Chenglong Zhang, Jingyao Cai, Xingwei Chen, Ningxin Su
G6D is a learning‑free, geometry‑driven RGB‑D 6D pose solver designed for robotic manipulation. It generates pose hypotheses via template‑based geometric matching and refines them using silhouette and depth consistency, requiring only an RGB‑D observation, an object mask, camera intrinsics, and a CAD model. The method offers adjustable accuracy‑computation trade‑offs, can run on CPU without GPUs, and has shown strong performance on LineMOD and BOP19 datasets, as well as in real‑world pick‑and‑place experiments.
By Yixuan Liang (Tsinghua University), William Chen (Sapient Intelligence), Yunan Wang (Tsinghua University), Jizhou Yan (Tsinghua University), Zhao Jin (Tsinghua University), Changling Liu (Sapient Intelligence), Chuxiong Hu (Tsinghua University)
The paper introduces the ACE framework, which uses spatially explicit inspection to guide embodied exploration. By combining evidence‑grounded perception with exposure‑informed movement, ACE provides a spatially resolved decision paradigm that improves cue assessment and movement direction. Experiments show ACE boosts navigation task success by 18.0% and exploration efficiency by 10.3% over previous baselines.
By Wenbin Wang, Xiang Bai, Yizhao Wang, Hang Sun, Dong Ren, Jie Qin, Qingquan Li, Bing Wang
The paper introduces ProAction, a multimodal dataset of 10,000 samples comprising visual, audio, and text inputs across 12 daily-life scenarios, designed to support the Proactive Robot Action Reasoning (ProRobo) problem. It presents a two-stage human-in-the-loop annotation pipeline that incorporates appraisal and Theory-of-Mind considerations to generate cognitively grounded high-level action labels. The authors benchmark multimodal large language models and propose MMC2Act, showing that training on ProAction significantly improves proactive action reasoning compared to general-purpose models.
By Zhihao Gu, Kechao Zhu, Yuanfeng Wu, Mohan Liu, Ankit Kumar Shaw, ChenDong Hong, Xuanyu Chen, Dengchen Mei, Xu Tianyi, Lin Wang
The paper introduces CSC, a calibrated-simplicity framework for detecting social bots in the era of large language models. CSC combines a simplified prototype-guided graph expert, calibrated simplex-constrained fusion, and a lightweight inconsistency expert to address modality conflict between semantic and structural signals. Experiments on TwiBot-22, TwiBot-20, and MGStBot-large demonstrate that CSC improves calibrated decision quality while maintaining competitive performance across benchmarks.
By Yipeng Qian, Pengjie Zhao, Chaoxi Niu
arXiv:2609.22299v1 Announce Type: cross
Abstract: When a robot faces unfamiliar physical conditions, a common approach is to collect evidence about what changed and adapt. For such diagnosis to impro...
By Zhengshu Zhang
arXiv:2609.22684v1 Announce Type: cross
Abstract: Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA...
By Wenzhuo Li, Qiongfeng Shi, Yi Zhou
arXiv:2609.23885v1 Announce Type: cross
Abstract: Robot foundation models learn manipulation from large demonstration corpora, but surgery is missing from those corpora: across the 780-hour Open-H su...
By Rhea Huang, David L. Matlock, Laurence Reich
arXiv:2609.24749v1 Announce Type: cross
Abstract: Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate wil...
By Shuaijun Liu, Chengyu Wu, Qifu Wen, Feiyang You, Chenglong Zhang, Shuyang Hao, Xi Lin, Ningxin Su
arXiv:2608.29601v2 Announce Type: replace-cross
Abstract: We present $N_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale mul...
By NeoteAI Team, Fudan TEAI Team
arXiv:2609.22090v1 Announce Type: new
Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present Ps...
By Joy Bose
arXiv:2609.22198v1 Announce Type: new
Abstract: The rapid adoption of large language models (LLMs) creates new opportunities for strategic content generation on online platforms, including potentiall...
By Valeria Lerman, Oren Rigbi, Yaniv Dover
arXiv:2609.22949v1 Announce Type: cross
Abstract: Existing prompt injection research focuses on single-model chatbot scenarios, where an attacker manipulates one LLM through crafted input. Multi-agen...
By Rudrendu Kumar Paul, Sourav Nandy