Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

Hugging Face Trending Papers
Sep 8

ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation

ReMoMask-2 is a retrieval‑augmented text‑to‑motion generation framework that improves on complex motion descriptions by addressing coarse retrieval and representation gaps. It introduces a structure‑aware RAG pipeline with Hierarchical Bidirectional Momentum contrastive learning, Semantic Spatial‑Temporal Attention, and Topology Structured Masking, and rebuilds the retrieval database in the generator’s latent space using a lightweight projector. Experiments on HumanML3D, KIT‑ML, and SnapMoGen show state‑of‑the‑art retrieval accuracy and the lowest FID scores, with a single mask‑transformer stage delivering faster inference than the previous two‑stage design.

Hugging Face Trending Papers
Sep 8

FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy

FPicker is a topology‑guided framework for filament tracing in low‑signal Cryo‑EM images. It combines a center‑endpoint representation with an open‑curve evolution module to model non‑cyclic connectivity, overcoming limitations of pixel‑wise segmenters, box‑based detectors, sequential trackers, and traditional active contours. On simulated benchmarks, FPicker improves mean spatio‑angular precision by over 40 % and reduces topological gap rates by more than 60 % under extreme noise, and it achieves state‑of‑the‑art performance on real EMPIAR data after fine‑tuning.

Hugging Face Trending Papers
Sep 8

3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

The paper introduces 3DWay, a method that predicts 3D consistent waypoints for robot manipulation using multi‑view images. By first generating 2D waypoints that are consistent across views and then triangulating them, the approach provides explicit 3D motion specifications while leveraging pretrained vision‑language models. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing the generalization of robot manipulation policies.

arXiv AI
Sep 7

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

The paper reviews how large language models have evolved into agents that can influence external environments through tool use, interface operation, delegation, state retention, virtual world inhabitation, and robotic control. It critiques the narrative of a single march toward autonomy, distinguishing model competence from system integration, persistence, and safe authority. The authors find that action-interface expansion is well documented, while robust completion, recovery, authorization, and independent verification remain less proven, and they propose a framework of justified delegation to guide future research.

By Linsen Zhu, Mengqing Cai
arXiv AI
Sep 7

How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions

The paper reports the first empirical study comparing how humans and large language models (LLMs) evaluate perceived moral agency (PMA) in both human and autonomous artificial agents within smart city scenarios. Using a validated PMA scale, 190 human participants and various LLMs were assessed, revealing that humans are perceived to have higher moral agency than artificial agents. When confronted with moral dilemmas, LLMs focus on situational factors such as harm severity and urgency, mirroring the context‑sensitivity observed in human raters.

By Fernanda Mansilla, Aloysius Tok, Bahia Guella\"i, Farah Benamara, Nancy F. Chen
arXiv AI
Sep 7

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.

By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
arXiv Computation and Language
Sep 7

Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models

The paper examines whether large language models (LLMs) exhibit conspiratorial tendencies, socio-demographic biases in this domain, and how easily they can be conditioned to adopt conspiratorial viewpoints. Using validated psychometric surveys, the authors find that LLMs partially align with conspiracy beliefs, that conditioning with demographic attributes yields uneven effects revealing latent biases, and that targeted prompts can readily shift responses toward conspiratorial stances. These findings underscore the vulnerability of LLMs to manipulation and the potential risks of deploying them in sensitive contexts.

By Francesco Corso, Francesco Pierri, Gianmarco De Francisci Morales
arXiv AI
Sep 7

Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution

The paper introduces a Cost-Aware Hierarchical Multi-Agent System (HMAS) for ransomware detection and family attribution that adaptively selects analysis modalities to balance accuracy and computational cost. Static analysis is used first, with dynamic and memory modalities added only when confidence is low or specialist agents disagree, guided by a cost model. Experiments show HMAS achieves high accuracy (96.57% binary detection, 0.90 macro‑F1 attribution) while reducing analysis cost by 43.97% and latency, with 56.05% of cases resolved using static evidence alone.

By Mubashar Iqbal, Asifullah Khan
arXiv Computer Vision
Sep 7

HiSfM: Disambiguating Structure-from-Motion via Scaffold-Anchored Hierarchical Reconstruction

HiSfM introduces a hierarchical coarse‑to‑fine Structure‑from‑Motion framework that enhances robustness and efficiency by building a scaffold of local communities and a compact skeleton using edge‑disjoint spanning trees. The method verifies skeletal edges with a two‑view disambiguator, constructs a stable scaffold as an anchor, and then registers remaining images for refinement. Experiments on ambiguity‑focused benchmarks and general datasets demonstrate that HiSfM avoids ambiguity‑induced failures, reduces runtime, and improves completeness compared to prior approaches.

By Ziding Zhao, Hainan Cui, Peilin Tao, Shuhan Shen
arXiv AI
Sep 7

CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review

The paper introduces CABAL, an end-to-end multi-agent simulation framework that models reviewer assignment in academic conferences using large language model-driven reviewer agents. It presents an affinity-guided collusive bidding strategy that forms collusion rings based on reviewer-paper affinities, leading to more effective target-paper capture and higher scores for colluding reviewers. Experiments show that while collusive bidding significantly increases target-paper capture and reviewer scores, overall conference-wide effects are modest, and existing bid-phase detectors offer limited detection capability.

By Jicheng Zhou, Kemou Li, Kahim Wong, Zheyuan Li, Zhuan Shi, Fengpeng Li, Haiwei Wu, Jiantao Zhou
arXiv AI
Sep 7

A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning

The paper presents a decentralized navigation framework for composite heterogeneous robots that integrates a large language model (LLM) policy agent, an Upper Confidence Bound (UCB) bandit, and a Double Deep Q-Network (Double DQN) controller. Each robot independently generates and refines policies at the round level using LLM inference, while the Double DQN handles tick-level action selection based on navigation variables and LLM priors. Across 30 rounds, the full configuration achieved all goals with the lowest median completion time (42 ticks) and a 25–39% improvement over other setups.

By Chongwen Dong, Mithun Paul Saint-Germain, Pinjari Asif, Carlo R. daCunha
arXiv Computer Vision
Sep 7

SocioGesture: Real-Time and Adaptive Social Gesture Perception for Human-Robot Interaction

SocioGesture is a real‑time, adaptive system for recognizing social gestures in human‑robot interaction. It employs a compact, confidence‑aware body‑hand skeleton representation and a lightweight dual‑stream model that fuses body motion with hand articulation, enabling low‑latency onboard recognition. The model is trained with occlusion‑aware skeleton corruption to handle missing hands, occluded arms, and unstable keypoints, and it can expand its gesture vocabulary during deployment by saving uncertain interaction segments for offline labeling.

By Wenjin Fu, Li-Fan Wu, Jerin Peter, Chip Huyen, Boyuan Chen, Jan Liphardt
arXiv Computer Vision
Sep 7

CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving

CoLMIN is an LLM-based framework for cooperative autonomous driving that addresses premature convergence to suboptimal solutions in multi-solution traffic scenarios. It introduces a Multi-Intent Negotiation module that generates multiple candidate driving intentions, an Evaluation-based Shallow Reflection Module that provides feedback to accelerate consensus, and a Deep Reflection Module that mitigates cognitive fixation by reflecting on negotiation histories. Experiments in the CARLA simulation show that CoLMIN outperforms existing methods in challenging interactive driving scenarios.

By Zhe Huang, Zhaoxin Fan, Shuo Wang, Wenjun Wu, Xuan Zhao, Min Liu
arXiv Computer Vision
Sep 7

Out-of-Distribution Semantic Occupancy Prediction

The paper introduces Out-of-Distribution Semantic Occupancy Prediction, a task that focuses on detecting unknown objects in 3D voxel space for autonomous driving. It proposes Realistic Anomaly Augmentation to create two new datasets, VAA-KITTI and VAA-KITTI-360, and presents the OccOoD framework, which uses Cross‑Space Semantic Refinement to improve OoD detection while maintaining semantic occupancy accuracy. Experiments show OccOoD achieves an AuROC of 65.50% and an AuPRCr of 31.83% within a 1.2 m radius, demonstrating strong generalization to real‑world urban scenes.

By Yuheng Zhang, Mengfei Duan, Kunyu Peng, Yuhang Wang, Ruiping Liu, Fei Teng, Kai Luo, Zhiyong Li, Kailun Yang
arXiv Computer Vision
Sep 7

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

The paper introduces MovingDroneCrowd++, a large-scale video dataset for dense crowd counting and tracking from moving drones, featuring varied flight altitudes, camera angles, and lighting. It presents two new methods: GD3A for Video Individual Counting and GIA-Track for Multi-Object Tracking, both leveraging group-wise density assignment and identity association to handle aerial challenges. Experiments demonstrate significant improvements, reducing counting error by 47.4% and boosting tracking accuracy by 64.6%.

By Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
arXiv Computer Vision
Sep 7

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.

By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
arXiv AI
Sep 7

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

MultihopSpatial is a new benchmark for Vision‑Language Models that focuses on multi‑hop, compositional spatial reasoning with queries ranging from 1 to 3 hops across varied spatial perspectives. It introduces the Acc@50IoU metric, which jointly evaluates answer selection and precise bounding‑box prediction, and provides a large‑scale training corpus, MultihopSpatial‑Train, to improve spatial intelligence. Evaluation of 37 state‑of‑the‑art VLMs shows that compositional spatial reasoning remains a significant challenge, and reinforcement learning fine‑tuning on the corpus boosts both intrinsic spatial reasoning and downstream embodied manipulation performance.

By Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang