arXiv Machine Learning By Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath

Trust Guided Decision Transformer

Read the original on arXiv Machine Learning →

The paper introduces Trust Guided Decision Transformer (TGDT), a method that mitigates performance degradation in Decision Transformers during long rollouts by monitoring the model’s next‑state prediction error. TGDT evaluates multiple recent context suffixes, filters out those whose prediction error exceeds a calibrated threshold, and then selects the highest‑value action from the remaining trusted suffixes using a frozen critic. Experiments on D4RL navigation and locomotion tasks demonstrate that TGDT reduces persistent high‑error runs and improves returns compared to vanilla Decision Transformer and other context‑control baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 21

Adaptive Rollout Truncation Based on Epistemic Uncertainty for Efficient Offline World Model Training

The paper introduces an adaptive rollout truncation method for offline world model training that uses epistemic uncertainty to decide when to stop autoregressive rollouts. By calibrating a threshold during a warm‑up phase, the approach replaces fixed‑horizon rollouts with uncertainty‑driven truncation, evaluated with ensemble and Monte Carlo dropout estimators. Experiments on ANYmal‑D and ANT demonstrate that this strategy matches or surpasses fixed‑horizon training while reducing cumulative rollout steps by about 72%.

By Nikodem Sebastian Zymla, Laurin Thiele, Johannes Pitz
arXiv AI
Aug 5

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies

arXiv:2608. 02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress.

By Inkyu Sa, Konstantin Stulov, Rajat Bhageria
arXiv Machine Learning
Sep 1

World Model Control by Trajectory Reachability Metrics

arXiv:2605.22164v2 Announce Type: replace Abstract: Latent world models can learn representations that contain information needed for control, while the downstream controller may still rank candidate...

By Liangyu Li, Shengzhi Wang, Libin Qiu, Mingliang Xiong, Qingwen Liu