arXiv:2608. 11658v1 Announce Type: cross Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive.
By Zijian Zhao, Sen Li
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
By Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim
arXiv:2606. 25526v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent.
By Bang Giang Le, Viet Cuong Ta
arXiv:2607. 19117v1 Announce Type: new Abstract: Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and continuous parameterization.
By Ubayd Ali Bapoo, Clement N Nyirenda
The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.
By Chuhao Qin, Evangelos Pournaras
arXiv:2406. 09770v2 Announce Type: replace-cross Abstract: Solving multi-objective optimization problems for large deep neural networks is a challenging task due to the complexity of the loss landscape and the expensive computational cost of training and evaluating models.
By Anke Tang, Li Shen, Yong Luo, Shiwei Liu, Han Hu, Bo Du, Dacheng Tao
The paper presents an autonomous agent that designs machine learning algorithms for wireless power control, eliminating manual specification of architecture, loss, and training details. Using an autoresearch protocol, the agent iteratively edits a training script, runs experiments, and evaluates changes against a single metric, ultimately achieving 99.5% of a reference solution with vastly reduced inference cost. The agent’s discovered output parameterization matches the exact max‑min‑optimal allocation at the minimum percentile for all trained weights, demonstrating a principled, scalable approach to a complex, NP‑hard problem.
By Ahmad Khan, Akram Bin Sediq, Sara Azadegi Naeini, Raviraj S. Adve
arXiv:2609.36770v1 Announce Type: new
Abstract: Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each...
By Aram Davtyan, Pablo Acuaviva, Sebastian Stapf, Paolo Favaro
arXiv:2608. 07532v1 Announce Type: new Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast.
By Mojtaba Eslami
arXiv:2608. 02391v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive.
By Zhiyuan Wang, Shengcai Liu, Jiahao Wu, Ning Lu, Hui Ouyang, Shaofeng Zhang, Haoze Lv, Ke Tang
arXiv:2609.14065v1 Announce Type: new
Abstract: When algorithmic predictions inform people's decisions, the models we deploy are performative and actively shape the data we see. This feedback loop be...
By Gabriele Farina, Juan Carlos Perdomo
arXiv:2506. 09046v3 Announce Type: replace-cross Abstract: Leveraging multiple Large Language Models (LLMs) has proven effective for addressing complex, high-dimensional tasks, but current approaches often rely on static, manually engineered multi-agent configurations.
By Xiaowen Ma, Yunpu Ma, Chenyang Lin, Sikuan Yan, Jinhe Bi, Zixuan Cao, Yijun Tian, Volker Tresp, Hinrich Schuetze