arXiv:2609.38156v1 Announce Type: new
Abstract: Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations ha...
By Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Shaoteng Liu, Lehan Yang, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen
arXiv:2609.37682v1 Announce Type: new
Abstract: The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the...
By Chu Zhang, Haoyu Jiang, Hongyuan Zhang, Hongbin Liu, Dong Yi
arXiv:2609.36359v1 Announce Type: new
Abstract: Graph-based approximate nearest neighbor search (ANNS) is widely used for large-scale semantic search. Its indices are constructed primarily based on g...
By Fangzhou Wu, Haike Xu, Sandeep Silwal
arXiv:2609.36828v1 Announce Type: new
Abstract: Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences wi...
By Wenxiao Fan, Jingling Fu, Lichen Ma, Yu He, Luohang Liu, Jinbao Xue, Ke Zhang, Junshi Huang, Kan Li
arXiv:2609.36835v1 Announce Type: new
Abstract: Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for...
By Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li, Adnan Aziz, Chunqiang Tang, Ang Li
arXiv:2609.37326v1 Announce Type: new
Abstract: On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller mode...
By Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar, Pascal Bouvry
arXiv:2609.36804v1 Announce Type: cross
Abstract: Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and gramm...
By Yitong Han, Nankai Lin, Juan Luo, Hongyan Wu, Lianxi Wang, Shengyi Jiang
arXiv:2609.36838v1 Announce Type: cross
Abstract: Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teache...
By Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu, Wei Li, Wen Luo, Yang Xu, Yufan Shen, Luke Mao, Yang Du, Asher Qin, Houfeng Wang
arXiv:2609.37712v1 Announce Type: cross
Abstract: Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, loca...
By GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui, Gaoyang Zhang, Jiangwei Xie, Jianshu Li, Jing Huang, Jingwen Bai, Mingqi Fang, Tao Fang, Weihong Zhang, Wenbo Du, Xiongfei Bai, Xuekang Zhu, Yinan Xia, Zhenming Wang, Jian Liu, Jingjing Liu, Xiang Qi, Weiqiang Wang
arXiv:2604.05468v3 Announce Type: replace
Abstract: Temporal knowledge graph (TKG) extrapolation is an important task that aims to predict future facts through historical interaction information with...
By Dongying Lin, Yinan Liu, Shengwei tang, Bin Wang, Xiaochun Yang
arXiv:2508.13009v5 Announce Type: replace
Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...
By Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou
arXiv:2609.37857v1 Announce Type: new
Abstract: Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. Howev...
By Zhenting Huang, Junnan Liu, Qianren Mao, Zhixing Tan, Bo Jiang
arXiv:2609.37988v1 Announce Type: new
Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This i...
By Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi
arXiv:2609.35800v1 Announce Type: new
Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composa...
By Nenad Banfic
arXiv:2609.35848v1 Announce Type: new
Abstract: On-board compression of synthetic aperture radar (SAR) phase history is bandwidth-critical, and block-adaptive quantization (BAQ) remains the operation...
By Alizishaan Khatri
arXiv:2609.35954v1 Announce Type: cross
Abstract: Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience i...
By Zhiwei Zhang, Huayu Deng, Fei Zhao, Jiayan Fu, Bin Liang, Kam-Fai Wong, Mu Chuan
arXiv:2609.36099v1 Announce Type: new
Abstract: Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce u...
By Aleksei Leonov, Nikita Kornilov, Zhenhe Zhang, Evgeny Burnaev, Iaroslav Koshelev, Alexander Korotin
arXiv:2609.36559v1 Announce Type: cross
Abstract: Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early...
By Yansong Liu, Rui Liu, Yuan Zuo, Hongwei Zhao, Da Fu, Fuwei Zhang, Fuzhen Zhuang, Yong Chen, Zhe Li
arXiv:2609.36638v1 Announce Type: new
Abstract: Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable i...
By Mingfeng Lin, Chengfei Cai, Lin Xu, Chengqian Ma, Yuxiang Wei, Liang Han
arXiv:2609.36642v1 Announce Type: new
Abstract: Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by lettin...
By Muyang Li, Jie Yang, Zhengyu Fang, Junchao Zhu, Zhengkun Xiao, Ruining Deng, Zhe Jiang, Shigang Chen