The paper proposes a CSI‑free hierarchical multi‑agent reinforcement learning framework for controlling reconfigurable reflective surfaces in millimeter‑wave networks. By replacing per‑element channel estimation with user localization data, the system uses a two‑tier neural architecture: a high‑level controller for discrete user‑to‑reflector assignments and low‑level controllers that optimize continuous focal points via MAPPO under a CTDE scheme. Deterministic ray‑tracing tests show RSSI gains of up to 7.79 dB over centralized PPO baselines and robust performance with sub‑meter localization errors for multiple users and reflector arrays.
By Hieu Le, Mostafa Ibrahim, Oguz Bedir, Jian Tao, Sabit Ekin
arXiv:2606. 26327v1 Announce Type: cross Abstract: In actor-critic reinforcement learning, network architectures are typically manually designed.
By Boyun Zhang, Chao Wang, Kai Wu
arXiv:2606. 20236v1 Announce Type: new Abstract: Many decision-making problems in computing and networking systems can be naturally formulated as cost-minimization problems under performance constraints.
By Federica Filippini
arXiv:2606. 10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems.
By Yavar Yeganeh, Mahsa Shekari, Nicla Frigerio, Daniele Pagano, Andrea Matta
arXiv:2606. 04328v1 Announce Type: cross Abstract: Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio resource management (RRM) toward artificial intelligence (AI)-driven RRM.
By Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci
arXiv:2607. 04758v1 Announce Type: new Abstract: Physical design quality-of-results~(QoR) optimization is hard and expensive.
By Shuo Ren, Zijin Cheng, Yaohui Han, Libo Shen, Leilei Jin, Wanting Tian, Rongliang Fu, Chao Wang, Bei Yu, Tsung-Yi Ho
The paper introduces a reinforcement‑learning‑guided evolutionary policy optimization framework for scheduling heterogeneous agile Earth observation satellites, addressing task selection, satellite assignment, and sequencing under diverse visibility windows, maneuvering constraints, energy use, and storage limits. It combines assignment‑based indirect encoding with decoder‑based cost evaluation to capture satellite‑dependent constraints while integrating task gain, energy savings, and load balance into a single utility metric. The resulting RLOSMEA algorithm uses reinforcement learning to select high‑level search operators, achieving higher weighted utility and more stable convergence than baseline metaheuristics across varied AEOS scenarios.
By He Wang, Junyu Wu, Hui Li, Yanjie Song, Witold Pedrycz, Liang Li
arXiv:2608.29490v1 Announce Type: cross
Abstract: Multi-agent systems in the real-world (e.g., drone swarms, autonomous cars, warehouse robots) must satisfy rich, temporal tasks while avoiding collis...
By Joe Eappen, Zikang Xiong, Shreyash S. Iyengar, Suresh Jagannathan
arXiv:2607. 25090v1 Announce Type: new Abstract: Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment interactions.
By Rushi Qiang, Changhao Li, Haotian Sun, Yuchen Zhuang, Chao Zhang, Bo Dai
arXiv:2602. 24115v2 Announce Type: replace Abstract: Open RAN (O-RAN) exposes rich control and telemetry interfaces across the Non-RT RIC, Near-RT RIC, and distributed units, but also makes it harder to operate multi-tenant, multi-objective RANs in a safe and auditable manner.
By Zhizhou He, Yang Luo, Xinkai Liu, Mahdi Boloursaz Mashhadi, Mohammad Shojafar, Merouane Debbah, Rahim Tafazolli
arXiv:2606. 00417v1 Announce Type: cross Abstract: To meet the stringent requirements of emerging applications and the increasingly complex network management and operation, the Next Generation Mobile Networks (NextG), or 6G, will adopt an AI-native architecture on the Core Network (CN).
By Maria Katarine Santana Barbosa, Kelvin L. Dias
arXiv:2606. 28339v1 Announce Type: cross Abstract: Industrial 6G networks require ultra-reliable, low-latency, and energy-efficient connectivity in dynamic and blockage-prone environments, where conventional terrestrial deployments often fail to ensure stable coverage.
By Marwan Dhuheir, Thang X. Vu, Symeon Chatzinotas