arXiv AI

Task-Specified Active Metrological Inspection with Measurement-Steered VLA Manipulation and Deterministic Evidence Gating

arXiv Computer Vision
Sep 18

INSPECT: Learning Robot View Selection from Assistant Use

INSPECT is a system that learns how a robot should choose its camera view during assembly inspection by observing a smart‑glasses assistant that answers part queries and guides the user. It uses techniques such as Presence‑Invariant TwinSwap for object evidence calibration, claim‑indexed supervision to separate evidence needs from camera changes, and object‑centered calibration to adapt view preferences to robot poses. In experiments on gearbox assemblies and angle‑grinder recordings, INSPECT outperforms other non‑oracle policies, improving view utility and decision accuracy.

By Di Wen, Kailun Yang, Wenhao Guo, Yitian Shi, Junwei Zheng, Yufan Chen, Ruiping Liu, Jiale Wei, Rania Rayyes, Kunyu Peng
arXiv AI
Aug 25

RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

The paper introduces RACO, a reliability‑aware adaptive coarse‑to‑fine navigation framework for inspection‑oriented UAV vision‑language navigation. It treats the coarse goal as a runtime hypothesis, using object‑level anchors to correct localization before and at the transition to the fine stage, and applies scale‑adaptive terminal refinement for near‑miss cases. RACO is evaluated on the new LG‑UVI inspection setting and outperforms the HETT baseline by 9.53 and 7.98 percentage points on validation‑unseen and test‑unseen, respectively, while improving inspection‑region arrival and reducing false verification risk.

By Sen Wang, Yiming Sun, Jiaxuan He, Pengfei Zhu
arXiv AI
Aug 18

DeepInsight II: One Trace from Benchmark to Robot

arXiv:2608. 16556v1 Announce Type: new Abstract: Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on which deployment actually turns remain fragmented across benchmark-specific simulators, embodiments, and interfaces.

By Siyi Li, Yuchen Kang, Wuliang Wang, Zhengjie Zhang, Jiangpin Liu, Jianhao Yao, Jie Chen
arXiv AI
Sep 3

Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework

The paper introduces TRACE, a Transparent Reasoning Architecture for Credible Execution, which provides an explainable AI-based decision framework for autonomous robots. TRACE structures decision-making into four auditable layers—Semantic Perception, Belief Reasoning, Action Synthesis, and Execution Verification—to ensure every action can be traced back to sensor evidence through documented causal chains. Experimental results on warehouse robot navigation show high evidence traceability (98.6%), temporal continuity (99.0%), and decision reconstructability (98.1%) across 500 simulated decision cycles.

By Cagri Temel
arXiv AI
Jun 12

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

arXiv:2606. 13578v1 Announce Type: cross Abstract: Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach.

By Baochang Ren, Xinjie Liu, Xi Chen, Yanshuo Liu, Chenxi Li, Daqi Gao, Zeqin Su, Jintao Xing, Zirui Xue, Rui Li, Xiangyu Zhao, Shuofei Qiao, Minting Pan, Wangmeng Zuo, Lei Bai, Dongzhan Zhou, Ningyu Zhang, Huajun Chen
arXiv AI
Sep 2

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.

By Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang
Hugging Face Trending Papers
Jun 1

Monitoring Agentic Systems Before They're Reliable

Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the failure landscape. At this maturity level, task-level error detection may be infeasible: structural failure modes mask the signal that task-level monitors are designed to detect.

arXiv AI
Sep 18

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

The paper introduces HALTER, a graph-based system that automates the reset and evaluation of long-horizon robot manipulation tasks. HALTER constructs a spatial scene graph from point clouds and vision models, uses an LLM to score rollouts, plan resets, and verify success, all without labeled success images. In experiments on a Franka arm, HALTER restores scenes in 76% of episodes, improves skill completion estimation, and reduces operator time by 72% compared to manual reset.

By Jing Jiang, Yue Yang, Xinkai Jiang, Gedas Bertasius, Daniel J. Szafir, Rudolf Lioutikov