arXiv Machine Learning By Bang Giang Le, Viet Cuong Ta

Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 25526v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning

The paper introduces QUATRO, a reinforcement‑learning approach for fine‑tuning large language models that enforces trust‑region constraints directly rather than relying on heuristic clipping. By deriving a principled objective, QUATRO provides explicit control over policy updates and stabilizes entropy during training. Experiments on mathematical reasoning benchmarks demonstrate that QUATRO maintains stable training even with higher learning rates and increased policy staleness.

By Doyeon Lee, Eunyi Lyou, Hyunsoo Cho, Sookyung Kim, Joonseok Lee, Jaemoo Choi