arXiv:2609.36687v1 Announce Type: cross
Abstract: Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and to...
By Jihwan Lee, Kleanthis Avramidis, Junhyeok Lee, Tiantian Feng, Najim Dehak, Shrikanth Narayanan
arXiv:2609.36982v1 Announce Type: cross
Abstract: Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners...
By Zhiwei Yang, Jiahua Yang, Huiru Lin, Xing Chen, Quanlong Guan
arXiv:2609.36995v1 Announce Type: cross
Abstract: Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-relate...
By Xingtong Ge, Yutong Wang, Lunjie Zhu, Haitao Lin, Fangyu Lin, Yushi Huang, Xin Zhang, Yi Zhang, Yu Liu, Jun Zhang
arXiv:2609.37001v1 Announce Type: cross
Abstract: Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the compu...
By Xingyu Jia, Baole Ai, Ang Wang, Kang Zhao, Yong Li
arXiv:2609.37170v1 Announce Type: cross
Abstract: Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation tr...
By Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang, Guangting Wang, Fengyun Rao, Jing Lyu, Dong Liu
arXiv:2609.37243v1 Announce Type: cross
Abstract: Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typica...
By Dae Ung Jo, Jongin Lim, YoungJoon Yoo, Daeho Um
arXiv:2609.37617v1 Announce Type: cross
Abstract: Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target....
By Yunzhe Li, Kyoungjun Park, Hongzi Zhu, Lili Qiu
arXiv:2609.37626v1 Announce Type: cross
Abstract: No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independe...
By Chuan Liu, Shuoming Zhang, Zhicheng Li, Qianqi Sun, Ruiyuan Xu, Qiuchu Yu, Xiyu Shi, Huimin Cui, Jiacheng Zhao
arXiv:2609.37905v1 Announce Type: cross
Abstract: Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical featur...
By Shivang Chopra, Fotis Iliopoulos, Zsolt Kira, Gaurav Menghani
arXiv:2609.38166v1 Announce Type: cross
Abstract: Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Atte...
By Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
arXiv:2609.31857v2 Announce Type: replace
Abstract: Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled....
By Junxuan Li, Arko Mukherjee, Soumyabrata Pal
arXiv:2609.33455v2 Announce Type: replace
Abstract: On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since e...
By Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo
arXiv:2609.34653v2 Announce Type: replace
Abstract: On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: t...
By Zongshang Shen, Wangsong Yin, Daliang Xu, Mengwei Xu, Xuanzhe Liu
arXiv:2512.20636v2 Announce Type: replace-cross
Abstract: Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppr...
By Dhananjay Saikumar, Blesson Varghese
arXiv:2602.03915v2 Announce Type: replace-cross
Abstract: Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be...
By Levi Lingsch, Georgios Kissas, Johannes Jakubik, Siddhartha Mishra
arXiv:2605.10335v2 Announce Type: replace-cross
Abstract: Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial mem...
By Yao Lu, Dengdong Fan, Shixun Zhang, Yonghong Tian
arXiv:2609.33923v2 Announce Type: replace-cross
Abstract: A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is wel...
By I Kennedy, T Kennedy
arXiv:2609.36222v1 Announce Type: new
Abstract: Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferrin...
By Ali Abbasi, Justin Shi, Soheil Kolouri
arXiv:2609.36484v1 Announce Type: new
Abstract: On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substan...
By Hao Li, MeiJia Chen, Weijie Ren, Donghan Li, Zijun Tian, Jingchun Huang, Naibo Wang
arXiv:2609.36584v1 Announce Type: new
Abstract: Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, s...
By Zhaoxian Wu, Tayfun Gokmen, Omobayode Fagbohungbe, T. Patrick Xiao, Tianyi Chen