arXiv Machine Learning

Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making

arXiv:2607. 18910v1 Announce Type: new Abstract: Sequential decision making in non-stationary and partially observable environments requires rapid adaptation to latent regime changes.

arXiv AI
Jun 2

Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning

arXiv:2606. 00780v1 Announce Type: cross Abstract: Offline meta-reinforcement learning leverages static datasets to enable agents to generalize to unseen environments by combining offline efficiency with meta-learning adaptability, yet it faces key challenges from context and policy distribution shifts.

By Fuyuan Qian, Menglong Zhang, Song Wang, Quanying Liu
arXiv AI
2d ago

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

The paper introduces Decision Titan, a variant of the Decision Transformer that incorporates Test‑Time Training (TTT) layers to store episodic memories in network parameters. It evaluates this architecture on the X‑Maze environment, showing that Decision Titan can learn long‑term dependencies up to 20 times longer than its context window and generalise to sequences 1.7 times longer than the training data. The study also finds that temporal generalisation depends on the choice of time embeddings and that the ability to learn long‑term dependencies hinges on how relevant information is encoded.

By Jude Waide, Robert Lieck
arXiv Computation and Language
3d ago

MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation

MetaSteer is a new method for steering large language models that learns nonlinear, context-dependent interventions applied to attention projection matrices. Unlike traditional linear, context-independent techniques, MetaSteer adapts its effects based on the input, requiring no linear concept-geometry assumption. Trained once on a pooled preference corpus, it transfers zero‑shot to unseen concepts and out‑of‑distribution contexts, matching or surpassing strong task‑specific baselines on multiple benchmarks and model families.

By Mehdi Jafari, Hao Xue, Flora Salim
arXiv Machine Learning
Sep 22

VISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking

arXiv:2609.23875v1 Announce Type: new Abstract: The rapid growth of resident space objects is increasing the complexity of space situational awareness sensor tasking, challenging classical optimizati...

By Miguel Leiva-V\'elez, Adalberto Claudio Quiros, Nicolas Gaston Rozado, Hodei Urrutxua, V\'ictor Rodr\'iguez-Fern\'andez
arXiv AI
Sep 25

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

IterSynth introduces a role-decoupled, iterative synthesis framework for deep search agents, separating planning and synthesis into distinct Planner and Synthesizer modules that maintain a persistent summary state. This design mitigates role coupling and context noise, while the new Role-Decoupled Policy Optimization (RDPO) enhances training by combining outcome rewards with turn-level rubric evaluations. Experiments on five long-horizon benchmarks show IterSynth-8B outperforming prior ≤8B agents by 4.2% and delivering significant zero-shot gains over ReAct on proprietary models.

By Xingyu Wu, Yuchen Yan, Zhengxi Lu, Siqi Chen, Xin ZHANG, Aiting Liu, Chao Deng, Jie Liu, Jin Ma, Jian Shao, Jun Xiao, Yongliang Shen
arXiv AI
Sep 2

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

ReNFT is a method that repairs mode collapse in diffusion generators after reward post‑training by internally recalibrating probability mass. It identifies suppressed alternatives through anti‑hub prompts and uses two policy‑dominated routes to generate counterfactual proposals, then applies reward‑based ranking and a joint‑and‑paired NFT update to restore diversity while preserving reward. Experiments on PickScore and GenEval show that ReNFT retains almost all of the original reward while significantly boosting diversity metrics.

By Yuchen Bao, Chao Wen, Haowei Wang, Ruoxin Chen, Donghao Luo, Jiahui Zhan, Wenjian Huang, Shen Chen, Yiting Wang, Taiping Yao, Chengjie Wang, Shouhong Ding, Jianguo Zhang