Train your first Decision Transformer
Related stories
Symmetry-Aware Transformer Training for Automated Planning
arXiv:2508. 07743v2 Announce Type: replace Abstract: While transformers excel in many settings, their application in the field of automated planning is limited.
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Differential Transformer V2
Vision Transformer Finetuning Benefits from Non-Smooth Components
arXiv:2602. 06883v3 Announce Type: replace Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness.
DiScoFormer: One transformer for density and score, across distributions
Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent
LiFT: Local Search via Linear Programming for Overfitting-Controlled Transformers
arXiv:2606. 16243v1 Announce Type: new Abstract: This paper proposes a Linear Programming (LP)-based local search framework for fine-tuning pretrained transformer models with explicit control against overfitting.
The State-Prediction Separation Hypothesis
arXiv:2607. 01218v1 Announce Type: cross Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions.
Stability of Transformers under Layer Normalization
arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
Transformer-based Encoder-Decoder Models
Training nGPT
arXiv:2608. 01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere.