arXiv:2609.23989v1 Announce Type: new
Abstract: Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of tr...
By Haixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang
arXiv:2609.22981v1 Announce Type: cross
Abstract: We present a dual-locking method for securing trained neural networks that combines key-driven index permutation with PIN-based watermarking based on...
By Iva Vasic, Jes\'us Mu\~noz-C\'adiz, Bata Vasic
arXiv:2609.23444v1 Announce Type: cross
Abstract: Engineering change order (ECO) is an important step in repairing timing and electrical violations during the late stages of chip design. Existing Age...
By Guoxiang Xu, Guozhen Ji, Zijian Luo, Zhengrui Chen, Qi Sun, Cheng Zhuo
arXiv:2609.24084v1 Announce Type: cross
Abstract: Open-weight large language models (LLMs) can be copied, modified, and redeployed behind black-box APIs, making post-release ownership verification di...
By Jiaxin Hong, Yuxin Peng, Hongyao Yu, Hao Fang, Shuoyang Sun, Bin Chen
arXiv:2609.24757v1 Announce Type: cross
Abstract: This paper presents the design, optimization, implementation, and on-board validation of a neural processing unit (NPU) accelerator for real-time veh...
By Daniel Gutierrez, Antonio Cuesta, Jorge Fe, Bruno Gutierrez, Rashed Al Koutayni
arXiv:2609.12551v2 Announce Type: replace-cross
Abstract: AI is beginning to make substantive contributions to LLM inference optimization. Existing AI optimizations are predominantly profiling-based....
By Ziyue Yang, Yuting Jiang, Lei Qu, Peng Cheng
arXiv:2609.25623v1 Announce Type: new
Abstract: More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a frozen copy of...
By Kanghui Tian, Siyuan Liu, Tianxiang Jiang, Shuai Dong, Yizhuo Li, Tian Ding, Yuan Guo, Songze Li, Haowen Hou, Congcong Wang, Yi Wang
arXiv:2609.25809v1 Announce Type: new
Abstract: Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly man...
By Yuanteng Chen, Qiwei Lai, Chen Tianqi, Peisong Wang, Yuantian Shao, Nanxin Zeng, Zhilei Liu, Chuangyi Li, Jing Liu, Jian Cheng
arXiv:2609.25963v1 Announce Type: new
Abstract: Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on h...
By Baher Mohammad, Ammar Ali, Stamatios Lefkimmiatis
arXiv:2609.26067v1 Announce Type: new
Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions, increasing flexibility but also parameter memory be...
By Kazi Ahmed Asif Fuad, Lizhong Chen
arXiv:2609.26333v1 Announce Type: new
Abstract: Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce me...
By Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh
arXiv:2609.26355v1 Announce Type: new
Abstract: Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted ma...
By Jiayan Fu, Hang Xu, Yong Zhang, Zhaokai Luo, Yao Hu, Dongyan Zhao, Mu Chuan
arXiv:2609.26708v1 Announce Type: new
Abstract: Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathema...
By Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng
arXiv:2609.25537v1 Announce Type: new
Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing laten...
By Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia, Yutaka Watanobe, Fang Liu
arXiv:2609.26368v1 Announce Type: new
Abstract: Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context dem...
By Jianyu Wei, Yizhao Gao, Qihao Zhang, Shimao Chen, Zhengju Tang, Yu Cheng, Shengjie Zhou, Zihan Jiang, Yifan Song, Hailin Zhang, Liang Zhao, Bo Yang, Gang Wang, Shijie Cao, Fuli Luo
arXiv:2609.25176v1 Announce Type: cross
Abstract: Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these...
By Lujia Bao, Qian Chen, Luyao Cheng, Chong Deng, Yuxiang Kong, Xiangang Li, Xu Li, Jiaqing Liu, Chao-Hong Tan, Haoyu Wang, Wen Wang, Xilou Wang, Junhao Xu, Liang Yi, Binbin Zhang, Qinglin Zhang, Qiquan Zhang
arXiv:2609.25498v1 Announce Type: cross
Abstract: Deploying Large Language Models for runtime operational triage incurs prohibitive latency (>100-500 ms), high VRAM requirements (>4-8 GB), and excess...
By Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i}
arXiv:2609.26073v1 Announce Type: new
Abstract: Pathology foundation models (PFMs) provide strong tile-level representations but remain difficult to interpret at the cellular and microenvironmental s...
By Yuxiang Xiao, Zhiwei Chen, Dan Dai, Wei Li, Tianyang Zhang, Yakun Ju, Yang Hu, Kaixiang Yang
arXiv:2609.26425v1 Announce Type: new
Abstract: KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficien...
By Jiaqi Zhao, Xiaobin Hu, Bo Yin, Junpeng Jiang, Miao Zhang, Shuicheng Yan
arXiv:2609.24526v2 Announce Type: replace
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints...
By Foundation Model, Li Auto Inc