arXiv:2504. 17768v3 Announce Type: replace-cross Abstract: Sparse attention offers a promising strategy to extend long-context capabilities in Transformer LLMs, yet its efficiency-accuracy trade-offs remain unclear due to the lack of comprehensive evaluation.
By Piotr Nawrot, Robert Li, Renjie Huang, Sebastian Ruder, Kelly Marchisio, Edoardo M. Ponti
The article investigates how Bayesian optimization can improve the ACTS parameter optimization suite for charged‑particle reconstruction. By comparing Expected Improvement and Upper Confidence Bound with TPE and random search on an eight‑parameter problem, extending the best method to fifteen parameters, and applying Expected Hypervolume Improvement for multi‑objective tuning, the study shows that Bayesian acquisition methods find strong configurations earlier and maintain advantages in held‑out validation. The results demonstrate that Bayesian optimization enhances ACTS auto‑tuning through more efficient evaluations, broader search spaces, and the ability to select from non‑dominated trade‑off solutions.
By Chance LaVoie, Qi Bin Lei, Rocky Bala Garg, Lauren Tompkins
Deploying Large Language Models (LLMs) in practice incurs substantial memory and computational costs. Post-training pruning (PTP) is an effective approach to reducing these costs by removing weights without additional training.
arXiv:2606. 01544v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) in practice incurs substantial memory and computational costs.
By Cheonjun Park
arXiv:2608. 06912v1 Announce Type: new Abstract: The top-$k$ operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and attention pruning.
By {\L}ukasz Struski, Joanna Wojciechowicz, Jakub Antczak, Marcin Mazur, Kamil Ksi\k{a}\.zek, Jacek Tabor
arXiv:2606. 21641v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that "warm-start" search from prior knowledge, proposing strong configurations in very few evaluations.
By Carson Rodrigues, Oysturn Vas, Isaiah Abner DCosta, Nithish Kumar Prabhakaran