CT‑SAFR is a multi‑layered verification framework designed to enhance the safety and faithfulness of Chain‑of‑Thought reasoning in autonomous robots. The framework achieves a 94.2% hallucination detection rate with sub‑500 ms latency, and a warehouse robot case study shows an 87% reduction in unsafe reasoning outputs. The study also offers recommendations for responsible deployment of reasoning‑capable autonomous robots.
By Cagri Temel
SAMV-DUSt3R is an end‑to‑end model that injects SAM2 2D masks into MV‑DUSt3R reconstruction to decouple objects from 3D scenes. A Cross Flow Mask Block steers the network toward the target instance, improving shape accuracy and achieving object‑level disentanglement without multi‑stage pipelines. A lightweight Spatial RankGNN selects the optimal reference view with 73.5% accuracy, and experiments show an 11% boost in average reconstruction precision over state‑of‑the‑art baselines.
By Langxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao, Tianhan Gao
The paper presents a probability‑wave framework for modeling the collective behavior of adaptive agents, deriving testable eigenmodes via a generalized behavioral intelligence (GBI) nonlocal probability‑wave equation. Empirical analysis of Chinese intraday stock market data shows that adaptive entangled game modes explain 89% of observed decision patterns, far exceeding predictions from neoclassical finance and supporting the Liu‑Chen‑Ao hypothesis of nonlocal entangled nerve fibers. The authors argue that incorporating these adaptive entangled game modules into AGI architectures can overcome limitations of conventional ANN‑based AI and enable more compact, efficient, and robust human‑like processing units for embodied intelligence and robotics.
By Haochen Li, Xinshuai Guo, Jingdong Ouyang, Wei Zhang, Leilei Shi
The paper introduces a taxonomy of seven distinct sources that contribute to the formation of Physical AI capabilities: Recorded-Experience, Predictive-Modeling, Evaluative-Interaction, Surrogate-Environment, Mechanism-Grounded, Embodied-Coupling, and Evolution-Driven. Through a systematic review of 49 evidence records, the authors demonstrate that all identified capabilities can be explained by combinations of these sources, achieving theoretical saturation within their defined scope. The framework aims to differentiate how capabilities arise from how they appear, facilitating analysis of explanation, transfer, replication, and governance aspects.
By Gang Chen
Grounded Continuation introduces a runtime verifier that classifies each utterance in an LLM conversation into one of eight epistemic operations and uses a symbolic engine to maintain a dependency map of claims and their supports. The verifier checks whether a new continuation is grounded by walking this map, a linear-time process that requires no additional LLM calls. On benchmarks such as ReviseQA and MemoryAgentBench, the verifier improves single-hop accuracy for several QA models, even enabling a 7B model to outperform GPT‑4o when guided by the verifier.
By Qisong He, Jinwei Hu, Xinmiao Huang, Changshun Wu, Yi Dong, Xiaowei Huang
The paper introduces TaMeSo‑bot, a soft‑wrist robotic system that uses tactile memory to perform robust object insertion tasks. It employs a Masked Tactile Trajectory Transformer (MAT³) to jointly model actions, tactile cues, force‑torque data, and proprioception, learning spatiotemporal representations through masked token prediction. Experiments on peg‑in‑hole tasks show that MAT³ outperforms baselines and adapts well to unseen pegs and conditions.
By Tatsuya Kamijo, Mai Nishimura, Nodoka Shibasaki, Jeremy Siburian, Cristian C. Beltran-Hernandez, Masashi Hamaya
ObstaDiff is a diffusion-policy framework that introduces a lightweight obstacle-aware visual encoder to generate structured representations of targets, obstacles, and background. By aligning these representations, the policy produces end-effector trajectories that focus on a target-centered bottleneck pose while accounting for surrounding obstacles. In real-robot greenhouse trials, ObstaDiff achieved a 75.41% task success rate and an 8.20% obstacle collision rate, outperforming existing imitation-learning baselines in cluttered agricultural settings.
By Jiawen Wang, Kevin Yao, Khalid Jawed
Wiggle and Go! is a two‑stage framework for zero‑shot rope manipulation that first performs a brief, safe wiggle action to infer rope parameters, then uses those parameters to condition a trajectory optimizer for goal‑conditioned execution. The method achieves 3.55 cm average accuracy on 3D target striking in real‑world tests, far outperforming uninformed baselines, and secures over 50% success on multi‑objective lobbing and draping tasks. Predicted parameters transfer well to unseen motions, with a 0.95 Pearson correlation between simulated and real rope dynamics, demonstrating task‑agnostic generalization without retraining.
By Arthur Jakobsson, Abhinav Mahajan, Karthik Pullalarevu, Krishna Suresh, Yunchao Yao, Yuemin Mao, Bardienus Duisterhof, Shahram Najam Syed, Jeffrey Ichnowski
The paper introduces SHAPER, a self‑evolving framework that enables embodied agents to adapt to new environments without retraining the underlying foundation model. SHAPER keeps model parameters frozen and instead evolves reusable skills and a context‑code harness through rollouts in the target environment, allowing the same model to act as both planner and optimizer. Experiments on VLABench and ESI‑Bench demonstrate that this skill‑and‑harness optimization outperforms pure execution, supervised fine‑tuning, and test‑time scaling baselines, showing that self‑evolving adaptation is viable when model training is costly or impractical.
By Peidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Yiqun Zhang, Zihan Wang, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.
By Vladislav Diuzhev, Dmitry Yudin
The paper introduces ORCH, a method that applies human organizational theory to create task‑specific hierarchical structures for large, heterogeneous embodied AI teams. Using ORCH, teams of up to 50 agents across 25 wildfire‑response missions outperformed four existing multi‑agent frameworks, achieving higher mission scores and greater execution efficiency. Both human‑designed and language‑model‑generated ORCH organizations improved performance, with hierarchical organization preserving concurrent activity while coordinating ordered transitions between mission phases.
By Zhengran Ji, Jonathan Hyun, Boyuan Chen
The paper explores soft prompting for few‑shot object detection with vision‑language models, showing that optimizing a small number of continuous prompt tokens—especially when placed at the cross‑modal boundary and initialized from an empty space token—can match LoRA performance while training far fewer parameters. Soft prompting also avoids catastrophic forgetting, transfers to newer models, and can be verbalized into readable prompts. The study extends these findings to manipulation tasks, indicating that VLMs already contain much of the necessary knowledge for specialized domains, and the main challenge is learning how to ask for it.
By Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti, Deva Ramanan
The article argues that embodied AI poses a significant governance lag, the delay between technological deployment and institutional response. It identifies three interlinked forms of lag—observational, institutional, and distributive—and proposes a compliance architecture featuring deployment visibility, stack-level accountability, trigger-based adjustments, and automatic distributional responses. The central policy challenge highlighted is ensuring governance systems become observable, responsive, and adaptive before disruption becomes entrenched.
By Shaoshan Liu
The paper introduces GLiBRL, a deep Bayesian reinforcement learning framework that uses generalized linear task models with learnable nonlinear basis functions. GLiBRL performs exact, sequential Bayesian inference without variational approximations, providing closed‑form posterior updates and a marginal likelihood. The method is permutation‑invariant, compatible with both off‑ and on‑policy algorithms, and achieves the best zero‑shot performance among eight meta‑RL baselines on MuJoCo locomotion and MetaWorld manipulation tasks.
By Jingyang You, Hanna Kurniawati
The paper introduces a densely labeled coastal dataset of over 1,000 images from Oahu, Hawaii, featuring 18 semantic classes and 7,400 annotated instances. Seven modern vision‑language models were evaluated using text‑to‑mask, mask‑to‑mask, and mask‑to‑text alignment experiments. Results show that broad landscape classes are recognized more accurately than conventional object and coastal classes, with coastal concepts posing the greatest challenge; however, performance differences are not solely due to environmental context, and alternative textual labels improve recognition of several coastal concepts.
By Seth Knoop, Chad R. Samuelson, Gabriel R. Slade, Brady Moon, Joshua G. Mangelson
The paper introduces ORCH, a method that applies human organization theory to create task‑specific hierarchical structures for large, heterogeneous embodied AI teams. By combining concurrent and sequential interdependence, ORCH outperforms four existing multi‑agent frameworks across 25 wildfire‑response missions, improving mission outcomes, execution efficiency, exploration, and computational resource use. Both human‑designed and language‑model‑generated ORCH organizations yield significant performance gains, with hierarchical organization enabling sustained concurrent activity and coordinated transitions in long‑horizon missions.
The paper introduces 2AM, a system that keeps task memory solely within a multimodal Agent while using a single RGB‑based, stateless Action Model to execute motions. By compiling interaction history into subtask language and optional 2D grasp/place/move hints, the Agent steers the Action Model, which is trained to tolerate imperfect guidance through dropout, noise, and jitter. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion without depth, geometry, or planners, vastly outperforming the best baseline.
The paper introduces EgoGenEval, a benchmark that assesses visual generators’ physical consistency during ego‑motion by measuring Camera Motion Grounding and Scene State Preservation across 1,400 cases and 2,360 target views. It shows that current pose‑free models struggle to maintain both camera motion and scene state, and that pairwise supervision in training does not simultaneously improve these aspects. The authors suggest a trajectory‑centric approach that couples self‑conditioned rollouts with explicit pose and visibility supervision to address these limitations.
arXiv:2609.08153v1 Announce Type: new
Abstract: Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, including but not limited to synthesis, re...
By Nian Wu, Nivetha Jayakumar, Jiarui Xing, Miaomiao Zhang
arXiv:2609.05516v1 Announce Type: cross
Abstract: Unified perception enables autonomous driving systems to perform object detection, drivable-area segmentation, and lane segmentation within a single...
By Zhiyuan Nie, Zixi Zhou, Xianbin Gu