arXiv:2605. 26660v2 Announce Type: replace Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in the ultra-low-bit regime remains challenging.
By Phong Nam Huu Nguyen, Khoi M. Le, Cong-Duy T Nguyen, Anh Tuan Luu, Thong Thanh Nguyen, Tho Quan
arXiv:2608. 07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2607. 06601v1 Announce Type: cross Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips whole transformer blocks, and KV-cache quantization compresses attention memory.
By Andrii Balashov, Olena Ponomarova
Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. Pruning and quantization address this, but rely on manual, expert choices and on algorithms that are hard to apply across architectures.
arXiv:2608. 05499v1 Announce Type: cross Abstract: Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices.
By Sadegh Jafari, Mohiuddin Bilwal, Fan Zhou, Brian Gelder, Ali Jannesari
arXiv:2602. 18109v3 Announce Type: replace Abstract: Real-time schedulers must reason about tight deadlines under strict compute budgets.
By Rong Fu, Yibo Meng, Zeyu Zhang, Ziming Guo, Jia Yee Tan, Xiaojing Du, Simon James Fong
The paper introduces AnySearch, a reinforcement‑learning framework that trains a single policy to perform budget‑aware search for large language models under any budget constraint. The training proceeds in two phases: first, the agent learns with explicit budget state injection and structured reasoning prompts under linearly decaying budgets; second, the scaffold is removed and the agent adapts to randomly sampled budgets that match deployment conditions. The reward combines answer accuracy and budget efficiency, with adaptive weighting to emphasize efficiency for high‑accuracy queries and reduce it for low‑accuracy ones. Experiments on seven QA benchmarks demonstrate that AnySearch outperforms baselines across all budget scales, generalizes to unseen constraints, and improves tool productivity without excessive token overhead.
By Xiaowei Sun, Jin Li, Yili Hong, Yikun Fu, Yanghua Xiao
arXiv:2107. 08183v2 Announce Type: replace Abstract: High-dimensional state and action spaces combined with sparse reward structures in reinforcement learning (RL) environments typically require advanced control architectures.
By JaeYoon Kim, Junyu Xuan, Christy Liang, Farookh Hussain
arXiv:2602. 03120v2 Announce Type: replace-cross Abstract: Post-Training Quantization (PTQ) is essential for deploying Large Language Models (LLMs) on memory-constrained devices, yet it renders models static and difficult to fine-tune.
By Yinggan Xu, Kajetan Schweighofer, Risto Miikkulainen, Xin Qiu
arXiv:2605. 26418v2 Announce Type: replace-cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every workload we test - so when, if ever, does DRL actually help?
By Guilin Zhang, Chuanyi Sun, Kai Zhao, Shahryar Sarkani, John Fossaceca
GAMMA is a post‑training framework that learns module‑wise precision preferences for mixed‑precision quantization of large language models. It optimizes a teacher‑forced hidden‑state reconstruction objective under an augmented Lagrangian constraint and then projects the learned preferences into exact budget‑feasible discrete assignments via integer programming. Because the learned preferences encode a stable sensitivity ranking, a single training run can be reused for any deployment budget, reducing per‑budget adaptation from hours to minutes and outperforming fixed‑precision baselines and search‑based methods on Llama and Qwen models.
By Zhangyang Yao, Haiyan Zhao, Haoyu Wang, Xu Han
arXiv:2605. 25054v2 Announce Type: replace-cross Abstract: Deploying deep neural networks on resource-constrained 6G edge devices demands aggressive compression with minimal accuracy loss.
By Ayush K. Varshney, Konstantinos Vandikas, \v{S}ar\=unas Girdzijauskas, Adam Orucu, Aneta Vulgarakis Feljan