Hugging Face Trending Papers

Destination Support Restoration for Finite-Set Multimodal Trajectory Prediction

arXiv Computer Vision
Sep 30

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

CST‑WM is a causally structured world model designed for embodied visual tracking, where a robot must keep a moving target visible and recover it after occlusion or drift. The model separates state into target‑evidence, robot, and observation branches, removing direct action‑to‑target‑evidence edges to prevent causal hallucination and instead letting actions influence evidence through robot motion and resulting views. Evaluated on EVT‑Bench, Habitat 3.0, and real‑world trials with a Unitree Go2 quadruped, CST‑WM outperforms reactive trackers and other world‑model baselines in following, distance control, safety, and re‑acquisition, achieving 20 of 30 successful real‑world recoveries versus 14 for TrackVLA.

By Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang
arXiv AI
Sep 18

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

The paper introduces ActObs, a supervised fine‑tuning method that, unlike standard approaches, also predicts environment observations in agent trajectories. While both ActObs and action‑only training perform similarly after initial fine‑tuning, ActObs diverges during subsequent reinforcement learning, yielding higher pass@k scores on several benchmarks and better cross‑domain task performance. The authors attribute this advantage to ActObs’s joint supervision, which preserves observation gradients and prevents the policy from over‑specializing on actions alone.

By Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan, Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah
arXiv Machine Learning
Jun 16

Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

arXiv:2606. 17043v1 Announce Type: cross Abstract: When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary outcome (success or failure), yet the actor update requires per-transition supervision.

By Tongyan Fang, Siyuan Huang, Naiyu Fang, Ganlong Zhao, Zhongjin Luo, Jianbo Liu, Xiaogang Wang, Ying Dong, Hongsheng Li
arXiv AI
Aug 20

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.

By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
arXiv AI
Sep 16

Learning-Guided Planning in Large Dynamic Action Spaces: Budgeted Tree Search for One-to-Many Mobile Charging

The paper introduces LP‑BTS, a learning‑guided planning framework for mobile charging in large, dynamic action spaces. It uses a graph proposal policy to narrow candidate stops, a value critic to evaluate leaf nodes, and edge‑budgeted PUCT to compare short simulated futures before action selection. Experiments on a 30‑scenario battery‑life benchmark show LP‑BTS achieving the highest survival and alive‑AUC, outperforming domain‑engineered baselines and heuristic policies.

By Liang-Ching Tao, Pi-Chung Wang
arXiv AI
Sep 4

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.

By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv Machine Learning
Aug 11

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.

By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
arXiv AI
Aug 5

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies

arXiv:2608. 02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress.

By Inkyu Sa, Konstantin Stulov, Rajat Bhageria