The paper studies how the design of Mixture-of-Experts (MoE) routers affects inference speed when combined with Speculative Decoding (SD). It shows that routers promoting high expert coactivation reduce memory transfer costs and improve runtime. By integrating a global load‑balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism, the authors achieve a 21% throughput gain over baseline MoEs while preserving accuracy.
By Kumari Nishu, Han-Byul Kim, Santosh Chilkunda, Maxwell Horton, Arnav Kundu, Mohammad Samragh, Lauren Hannah, Mohammad Sekhavat, Nikhil Bhendawade, Manuel Ciosici, Iman Mirzadeh, Keivan Alizadeh Vahid, David Harrison, Irina Belousova, Mehrdad Farajtabar, Minsik Cho
The paper introduces a new compression pipeline for scientific simulation data that combines Residual Vector Quantization (RVQ) with a U‑Net post‑processing network to correct pixel‑space residuals, followed by a Guaranteed Autoencoder (GAE) that enforces block‑wise error bounds. The U‑Net is trained to predict spatially structured residuals, addressing limitations of latent‑space only approaches. Experiments on S3D, JHTDB, and E3SM datasets show improved NRMSE and compression ratios compared to RVQ alone while maintaining strict error guarantees.
By Surya Majumder, Liangji Zhu, Sanjay Ranka, Anand Rangarajan
arXiv:2609.23686v1 Announce Type: new
Abstract: Patch-based autoregressive time-series forecasting often ties input representation, learned transitions, and recursive execution to one patch length. W...
By Ziang Li, Yue Huang, Guoxu Zhou, Na Han, Jie Wen, Lunke Fei, Xiaozhao Fang
arXiv:2609.23900v1 Announce Type: new
Abstract: Tree speculative decoding verifies multiple candidate continuations in one target forward pass. For attention-only transformers, the verifier mainly ne...
By Zhiyuan Ma
arXiv:2609.24209v1 Announce Type: new
Abstract: The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent wor...
By Chenming Shang, Yujin Tang, Jun Jie Ou Yang, Ruize Xu, Adam Breuer, Nikhil Singh
arXiv:2609.24401v1 Announce Type: new
Abstract: Structured pruning is a model compression technique that is used to reduce the computational cost of deploying deep neural networks on resource-constra...
By Mindula Illeperuma, Rafael Pina, Charuka Herath, Sharmarke A. Gabayre, Varuna De Silva
arXiv:2609.24432v1 Announce Type: new
Abstract: Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teache...
By Huanxin Sheng, Zhiling Ye, Haonan Wang, Jian Wang, Jinjie Gu, Jian Kang
arXiv:2609.24464v1 Announce Type: new
Abstract: Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial pr...
By Armin Oliya, Aleksandra Sawczuk, Rados{\l}aw Bia{\l}obrzeski
arXiv:2609.24947v1 Announce Type: new
Abstract: Neural operators evaluate parametric partial differential equations cheaply but degrade sharply outside their training distribution. Physics-informed n...
By S. Mohammad Mousavi, Teeratorn Kadeethum, Nikolaos Bouklas, Somdatta Goswami
arXiv:2609.22215v1 Announce Type: cross
Abstract: Knowledge distillation can transmit unintended behavioral traits from a teacher model to a student through training data that appear semantically unr...
By Atsushi Yanagisawa, Brendan Gho, Rajendran Ramesh Babu Manoj Narender, Kevin Zhu, Madhur Panwar, Antonio Mari
arXiv:2609.23038v1 Announce Type: cross
Abstract: Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requi...
By Kaixiang Yao, Xu Wang, Miao Pan, Hu Xiyue, Weishi Wang, Daniel Dahlmeier, Jintao Chen, Yongliang Shen, Xuhong Zhang, Wenqi Zhang
arXiv:2609.23340v1 Announce Type: cross
Abstract: Pitch estimation on an edge device is constrained in three ways at once. The model must be small, it must stay accurate when the input is noisy, and...
By Venkat Suprabath Bitra, Homayoon Beigi
arXiv:2609.23376v1 Announce Type: cross
Abstract: Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sourc...
By Yi Xu, Cheng Chen, Wenzhuo Lei
arXiv:2609.23377v1 Announce Type: cross
Abstract: Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can b...
By Jie Zhao, Ziyu Jiang, Suhang Zheng, Minghui Shan, Xiaoxiao Xu, Lin Qu
arXiv:2609.23668v1 Announce Type: cross
Abstract: Inferring network topology from noisy node observations is a central problem in graph signal processing. In this paper, we consider Laplacian-constra...
By Christoffer Kjellson, Claudio Altafini, Emma Tegling
arXiv:2609.23703v1 Announce Type: cross
Abstract: Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrat...
By Kemal Kirtac
arXiv:2609.24379v1 Announce Type: cross
Abstract: Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entan...
By Gautam Ranka, Shubham Santosh Pandere, Aiden Dsouza
arXiv:2609.24569v1 Announce Type: cross
Abstract: Over the past decade, a growing body of research has shown that $\gamma$-weak submodularity broadly arises in numerous subset selection tasks, includ...
By Shi Fu, Youming Qiao, Dacheng Tao, Zongqi Wan, Qixin Zhang
arXiv:2502.01562v3 Announce Type: replace
Abstract: As the general capabilities of artificial intelligence (AI) agents continue to evolve, their ability to learn to master multiple complex tasks thro...
By Minttu Alakuijala, Ya Gao, Georgy Ananov, Samuel Kaski, Pekka Marttinen, Alexander Ilin, Harri Valpola
arXiv:2505.19893v2 Announce Type: replace
Abstract: Large language model pretraining is compute-intensive, yet many tokens contribute marginally to learning, resulting in inefficiency. We introduce E...
By Melis Ilayda Bal, Volkan Cevher, Michael Muehlebach