arXiv:2507. 10142v2 Announce Type: replace Abstract: Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the assumptions under which algorithms are designed and evaluated.
By Siyi Hu, Mohamad A Hady, Jianglin Qiao, Jimmy Cao, Mahardhika Pratama, Ryszard Kowalczyk
The paper introduces Transfer Learning for Evolving Domains (TrED), a framework that models how data availability changes over time in real-world applications. TrED treats the entire trajectory of model updates as a single learning problem, rather than isolated snapshots, and defines a data availability process, a flexible learning protocol, and an evaluation criterion that scores the whole trajectory. The authors review existing transfer learning methods, noting that most are tailored to specific regimes and do not optimize the full trajectory, and argue that TrED is a well‑posed, unsolved research direction.
By Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Pedro Ribeiro, Pedro Saleiro, Pedro Bizarro, Carlos Soares
arXiv:2606. 24601v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives.
By Anurag Akula, Satheesh K. Perepu, Abhishek Sarkar, Kaushik Dey
arXiv:2608. 11658v1 Announce Type: cross Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive.
By Zijian Zhao, Sen Li
arXiv:2606. 09825v1 Announce Type: cross Abstract: Training reinforcement learning (RL) policies from scratch is costly: it requires careful reward and environment design, extensive tuning, and substantial computation.
By Anton Bolychev, Georgiy Malaniya, Sinan Ibrahim, Pavel Osinenko
arXiv:2607. 25663v1 Announce Type: new Abstract: Transformer adaptation is typically distributed across model depth, even when the intended change is narrow.
By Rebecca Ramnauth, Brian Scassellati