Hugging Face Blog

Train your first Decision Transformer

arXiv Machine Learning
Sep 29

Trust Guided Decision Transformer

The paper introduces Trust Guided Decision Transformer (TGDT), a method that mitigates performance degradation in Decision Transformers during long rollouts by monitoring the model’s next‑state prediction error. TGDT evaluates multiple recent context suffixes, filters out those whose prediction error exceeds a calibrated threshold, and then selects the highest‑value action from the remaining trusted suffixes using a frozen critic. Experiments on D4RL navigation and locomotion tasks demonstrate that TGDT reduces persistent high‑error runs and improves returns compared to vanilla Decision Transformer and other context‑control baselines.

By Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath
arXiv Computation and Language
Sep 7

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

The paper introduces an influence score that measures how much each attention head contributes to classification decisions in Transformer models, specifically for prompt injection detection. The score blends directional effects on logits with structural impact within the residual stream, allowing analysis at head, layer, and network scales. When applied to a DeBERTa model, the framework uncovers different decision patterns for correct versus incorrect predictions, offering a balanced approach between detailed circuit analysis and global output methods.

By Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi
arXiv AI
Jul 2

The State-Prediction Separation Hypothesis

arXiv:2607. 01218v1 Announce Type: cross Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions.

By Giovanni Monea, Nathan Godey, Kiant\'e Brantley, Yoav Artzi