The paper proposes a developmental framework for autonomous artificial agents that emphasizes learning social norms and alignment through direct interaction with dynamic environments. It argues that intrinsic motivations such as curiosity and competence can guide exploration, but also complicate alignment with human goals. By drawing parallels to child development, the authors suggest that regulatory sandboxes serve as pedagogical spaces where agents gradually acquire moral agency and adapt their behaviors through experience and cooperation.
By Marica Notte, Ludovica Marinucci, Vieri Giuliano Santucci
The paper introduces 2AM, a system that separates memory and action execution in long‑horizon robot manipulation. 2AM stores task memory exclusively in a multimodal Agent, while a single RGB‑based, stateless Action Model performs motion based on language and optional 2D hints. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion—over 61 points higher than the best baseline—demonstrating that agent‑side memory and precise steering of the Action Model can substantially improve performance.
By Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry
The paper introduces **terms.txt**, a new protocol that extends the traditional robots.txt format to allow website owners to specify per‑path, per‑purpose access rules for automated agents. It proposes an origin‑enforced exchange using Web Bot Auth signatures, signed intent, delegation tokens, HTTP 402 negotiation, and signed receipts to enforce, audit, and contract these rules. A lightweight implementation adds only 0.20 to 0.65 ms per request on a single vCPU.
By Rajarshi Chowdhury
The paper introduces an admission‑audit protocol for continual embodied agents, arguing that update admission should balance error control with retained learning opportunities within a fixed interaction budget. It critiques a range‑based confidence gate for failing to certify unchanged old‑task behavior, and proposes a paired‑binomial construction that reduces this burden when outcome disagreements are rare. Experiments on a one‑step pushing diagnostic show that fresh paired checks admit a significant portion of updates while the range‑based gate admits none, and a learned‑dynamics stress test helps distinguish model bias from feedback‑selection error.
By Qinzhen Ma, Ruihai Wu
The paper introduces a pragmatic information theory that unifies communication, control, and decision-making through the isoteleia mapping, which formalizes equifinality by treating distinct semantic paths that lead to the same optimal action as pragmatically equivalent. It establishes a three-tier hierarchy of syntactic, semantic, and pragmatic information, defines pragmatic entropy, mutual information, channel capacity, and rate-distortion, and proves coding theorems that generalize Shannon’s results. The authors also present pragmatic value and cost of information, a Lagrangian dual framework for cross-layer optimization, and a pragmatic efficiency bound that quantifies the maximum net utility for resource-constrained intelligent systems, extending the theory to continuous messages and dynamic settings.
By Kai Niu, Ping Zhang
The paper introduces a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. It combines a fixed phase-dependent reflex controller with a reinforcement learning policy that adjusts reflex gains and thresholds for hip swing, knee support, and ankle propulsion. Experiments show the method produces physiologically plausible walking with improved kinematic accuracy, dynamic consistency, symmetry, and stride consistency, and remains robust to muscle weakness and perturbations without retraining.
By Jian Zhou, Xingyu Zhang, Rui Ma, Yu Cao, Shane Xie, Zhi-qiang Zhang
The paper introduces EgoGenEval, a new benchmark that assesses the physical consistency of visual generators under ego‑motion by measuring Camera Motion Grounding and Scene State Preservation across 1,400 cases and 2,360 target views. Experiments on 16 pose‑free generators and two pose‑conditioned references show that current models struggle to maintain both camera motion and scene state simultaneously. A follow‑up study using EgoGen‑Train demonstrates that pairwise supervision does not effectively improve both metrics together, suggesting the need for a trajectory‑centric training paradigm.
By Yilin Long, Chenming Zhu, Zitang Gou, Jingli Lin, Tai Wang
The paper introduces GLiBRL, a deep Bayesian reinforcement learning framework that uses generalized linear task models with learnable nonlinear basis functions. GLiBRL performs exact, sequential Bayesian inference without variational approximations, providing closed‑form posterior updates and a marginal likelihood. The method is permutation‑invariant, compatible with both off‑ and on‑policy algorithms, and achieves the best zero‑shot performance among eight meta‑RL baselines on MuJoCo locomotion and MetaWorld manipulation tasks.
By Jingyang You, Hanna Kurniawati
The paper introduces a densely labeled coastal dataset of over 1,000 images from Oahu, Hawaii, featuring 18 semantic classes and 7,400 annotated instances. Seven modern vision‑language models were evaluated using text‑to‑mask, mask‑to‑mask, and mask‑to‑text alignment experiments. Results show that broad landscape classes are recognized more accurately than conventional object and coastal classes, with coastal concepts posing the greatest challenge; however, performance differences are not solely due to environmental context, and alternative textual labels improve recognition of several coastal concepts.
By Seth Knoop, Chad R. Samuelson, Gabriel R. Slade, Brady Moon, Joshua G. Mangelson
The paper introduces SHAPER, a self‑evolving framework that enables embodied agents to adapt to new environments without retraining the underlying foundation model. SHAPER keeps model parameters frozen and instead evolves reusable skills and a context‑code harness through rollouts in the target environment, allowing the same model to act as both planner and optimizer. Experiments on VLABench and ESI‑Bench demonstrate that this skill‑and‑harness optimization outperforms pure execution, supervised fine‑tuning, and test‑time scaling baselines, showing that self‑evolving adaptation is viable when model training is costly or impractical.
By Peidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Yiqun Zhang, Zihan Wang, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
The paper explores soft prompting for few‑shot object detection with vision‑language models, showing that optimizing a small number of continuous prompt tokens—especially when placed at the cross‑modal boundary and initialized from an empty space token—can match LoRA performance while training far fewer parameters. Soft prompting also avoids catastrophic forgetting, transfers to newer models, and can be verbalized into readable prompts. The study extends these findings to manipulation tasks, indicating that VLMs already contain much of the necessary knowledge for specialized domains, and the main challenge is learning how to ask for it.
By Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti, Deva Ramanan
The paper investigates how inaccuracies in pretrained occupancy networks affect active mapping robots that select camera viewpoints to reconstruct unknown 3D scenes. By fixing the planner and varying the occupancy representation—ranging from no completion to ground‑truth occupancy—the authors find that correcting false positives or false negatives alone does not reliably improve coverage, highlighting a disconnect between occupancy accuracy and planning performance. They propose a dynamic filtering strategy that retains predictions in unexplored space while suppressing unsupported occupancy based on online observations, which preliminarily shows it can steer viewpoint selection toward reachable surfaces that would otherwise remain unseen.
By Jiahui Zhang, Gongbo Liang, Yu Zhang
The article argues that embodied AI poses a significant governance lag, the delay between technological deployment and institutional response. It identifies three interlinked forms of lag—observational, institutional, and distributive—and proposes a compliance architecture featuring deployment visibility, stack-level accountability, trigger-based adjustments, and automatic distributional responses. The central policy challenge highlighted is ensuring governance systems become observable, responsive, and adaptive before disruption becomes entrenched.
By Shaoshan Liu
The paper presents a 652,157‑parameter action‑conditioned visuotactile world model designed for lifting tasks, integrating behavior cloning, policy learning in imagination, reactive implicit Q‑learning, and model‑assisted force feedback. Experiments on 120 fresh MuJoCo environments and additional ID environments show that visuotactile dynamics reduce force‑action‑effect mean absolute error from 0.413 N to 0.338 N, and model‑assisted feedback boosts force‑budgeted success from 73.3 % to 93.3 %. Imagined reinforcement learning achieves 11.9 % pooled joint success compared to 25.0 % for reactive IQL, with further stress testing adding 330 executions.
By Qinzhen Ma (Rice University)
MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.
By Vladislav Diuzhev, Dmitry Yudin
ObstaDiff is a diffusion-policy framework that introduces a lightweight obstacle-aware visual encoder to generate structured representations of targets, obstacles, and background. By aligning these representations, the policy produces end-effector trajectories that focus on a target-centered bottleneck pose while accounting for surrounding obstacles. In real-robot greenhouse trials, ObstaDiff achieved a 75.41% task success rate and an 8.20% obstacle collision rate, outperforming existing imitation-learning baselines in cluttered agricultural settings.
By Jiawen Wang, Kevin Yao, Khalid Jawed
The paper introduces HuRo, a dataset of 630K robotized episodes derived from diverse human videos, created via a pipeline that aligns observations and actions for robotic use. Experiments on four real‑world manipulation tasks show that scaling robotized pretraining boosts task completion from 51.5% to 80.3% and improves out‑of‑distribution performance under spatial and visual shifts. Ablation studies reveal that visual robotization enhances robustness and that end‑to‑end pretraining with retargeted actions outperforms visual‑only transfer.
By Jinho Jeong, Se June Joo, Jaehyun Kang, Dongyun Kim, Yena Kim, Hanjung Kim, Seon Joo Kim
Valerant is a training‑free framework that turns a pretrained action‑conditioned world model into a world action model (WAM) capable of exploring and building 3D game maps. By combining predictive visual rollouts with SLAM‑based spatial reconstruction and exploration‑driven action selection, it converts a single image into a persistent, navigable 3D geometry. This approach extends WAM‑based interaction beyond 2D visual simulation and reduces manual effort in 3D game‑map creation.
By Yiran Qiao, Feng Wang, Jing Ma
SAMV-DUSt3R is an end‑to‑end model that injects SAM2 2D masks into MV‑DUSt3R reconstruction to decouple objects from 3D scenes. A Cross Flow Mask Block steers the network toward the target instance, improving shape accuracy and achieving object‑level disentanglement without multi‑stage pipelines. A lightweight Spatial RankGNN selects the optimal reference view with 73.5% accuracy, and experiments show an 11% boost in average reconstruction precision over state‑of‑the‑art baselines.
By Langxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao, Tianhan Gao
Efficient Diversity-based Experience Replay (EDER) is introduced to enhance learning efficiency in deep reinforcement learning. It uses a determinantal point process to model sample diversity, prioritizes replay based on this diversity, and incorporates Cholesky decomposition and rejection sampling to handle large state spaces. Experiments on MuJoCo, Atari, and Habitat show that EDER significantly improves learning efficiency and performance in high‑dimensional, realistic environments.
By Kaiyan Zhao, Yiming Wang, Yuyang Chen, Yan Li, Leong Hou U, Xiaoguang Niu