Hugging Face Trending Papers

A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization

arXiv Machine Learning
Sep 25

A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization

This paper proposes a particle‑swarm‑assisted gradient meta‑learning (PSA‑GML) algorithm to jointly optimize the transmit precoder and the transmission/reflection coefficients of a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR‑RIS) for maximizing weighted sum rate in a multi‑user downlink. The method first transforms the non‑convex problem via amplitude‑split parameterization and collapsed precoder representation, then uses particle swarm optimization to generate a robust warm start for the STAR‑RIS coefficients, and finally refines both coefficients and precoder with a coordinate‑wise LSTM meta‑optimizer trained by first‑order gradient meta‑learning. Numerical results demonstrate that PSA‑GML achieves an 11.06 bits/s/Hz weighted sum rate at 10 dB, outperforming conventional alternating optimization by 13.1 % and the random‑phase scheme by 35.1 %, while also showing strong zero‑shot transfer across regimes.

By Kang Zhou
arXiv AI
Sep 10

Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors

The paper proposes a CSI‑free hierarchical multi‑agent reinforcement learning framework for controlling reconfigurable reflective surfaces in millimeter‑wave networks. By replacing per‑element channel estimation with user localization data, the system uses a two‑tier neural architecture: a high‑level controller for discrete user‑to‑reflector assignments and low‑level controllers that optimize continuous focal points via MAPPO under a CTDE scheme. Deterministic ray‑tracing tests show RSSI gains of up to 7.79 dB over centralized PPO baselines and robust performance with sub‑meter localization errors for multiple users and reflector arrays.

By Hieu Le, Mostafa Ibrahim, Oguz Bedir, Jian Tao, Sabit Ekin
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv AI
Jun 4

Generalizable Multi-Task Learning for Wireless Networks Using Prompt Decision Transformers

arXiv:2606. 04328v1 Announce Type: cross Abstract: Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio resource management (RRM) toward artificial intelligence (AI)-driven RRM.

By Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci
arXiv AI
Jul 22

ISO: An RLVR-Native Optimization Stack

arXiv:2607. 19331v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood.

By Hanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David Gonz\'alez-Mart\'inez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang