arXiv Machine Learning By Yuyang Shen, Shan Dai, Daimin Chen

Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making

Read the original on arXiv Machine Learning →

arXiv:2607. 18910v1 Announce Type: new Abstract: Sequential decision making in non-stationary and partially observable environments requires rapid adaptation to latent regime changes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 2

Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning

arXiv:2606. 00780v1 Announce Type: cross Abstract: Offline meta-reinforcement learning leverages static datasets to enable agents to generalize to unseen environments by combining offline efficiency with meta-learning adaptability, yet it faces key challenges from context and policy distribution shifts.

By Fuyuan Qian, Menglong Zhang, Song Wang, Quanying Liu
arXiv AI
2d ago

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

The paper introduces Decision Titan, a variant of the Decision Transformer that incorporates Test‑Time Training (TTT) layers to store episodic memories in network parameters. It evaluates this architecture on the X‑Maze environment, showing that Decision Titan can learn long‑term dependencies up to 20 times longer than its context window and generalise to sequences 1.7 times longer than the training data. The study also finds that temporal generalisation depends on the choice of time embeddings and that the ability to learn long‑term dependencies hinges on how relevant information is encoded.

By Jude Waide, Robert Lieck
arXiv Computation and Language
3d ago

MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation

MetaSteer is a new method for steering large language models that learns nonlinear, context-dependent interventions applied to attention projection matrices. Unlike traditional linear, context-independent techniques, MetaSteer adapts its effects based on the input, requiring no linear concept-geometry assumption. Trained once on a pooled preference corpus, it transfers zero‑shot to unseen concepts and out‑of‑distribution contexts, matching or surpassing strong task‑specific baselines on multiple benchmarks and model families.

By Mehdi Jafari, Hao Xue, Flora Salim