Train your first Decision Transformer
Related stories
Symmetry-Aware Transformer Training for Automated Planning
arXiv:2508. 07743v2 Announce Type: replace Abstract: While transformers excel in many settings, their application in the field of automated planning is limited.
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Differential Transformer V2
Trust Guided Decision Transformer
The paper introduces Trust Guided Decision Transformer (TGDT), a method that mitigates performance degradation in Decision Transformers during long rollouts by monitoring the model’s next‑state prediction error. TGDT evaluates multiple recent context suffixes, filters out those whose prediction error exceeds a calibrated threshold, and then selects the highest‑value action from the remaining trusted suffixes using a frozen critic. Experiments on D4RL navigation and locomotion tasks demonstrate that TGDT reduces persistent high‑error runs and improves returns compared to vanilla Decision Transformer and other context‑control baselines.
Vision Transformer Finetuning Benefits from Non-Smooth Components
arXiv:2602. 06883v3 Announce Type: replace Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness.
DiScoFormer: One transformer for density and score, across distributions
Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent
Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
The paper introduces an influence score that measures how much each attention head contributes to classification decisions in Transformer models, specifically for prompt injection detection. The score blends directional effects on logits with structural impact within the residual stream, allowing analysis at head, layer, and network scales. When applied to a DeBERTa model, the framework uncovers different decision patterns for correct versus incorrect predictions, offering a balanced approach between detailed circuit analysis and global output methods.
Diffusion-Inspired Reconfiguration of Transformers for Uncertainty Calibration
arXiv:2602.08920v3 Announce Type: replace Abstract: Uncertainty calibration in pre-trained transformers is critical for their reliable deployment in risk-sensitive applications. Yet, most existing pr...
LiFT: Local Search via Linear Programming for Overfitting-Controlled Transformers
arXiv:2606. 16243v1 Announce Type: new Abstract: This paper proposes a Linear Programming (LP)-based local search framework for fine-tuning pretrained transformer models with explicit control against overfitting.
The State-Prediction Separation Hypothesis
arXiv:2607. 01218v1 Announce Type: cross Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions.