arXiv:2609.39405v1 Announce Type: new
Abstract: Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced b...
By Jingang Zhou, Feiyu Han, Han Zhu, Yuyi Zhou, Ruiyang Zhang, Jian Xu, Sirui Gao, Qingpei Guo, Xu-Yao Zhang
arXiv:2609.39489v1 Announce Type: new
Abstract: Sample-level reliability heterogeneity is common in deep time series learning. Standard training pipelines apply a uniform regularization setting to al...
By Siru Zhong, Senzhang Wang, James T. Kwok, Yuxuan Liang
arXiv:2609.39816v1 Announce Type: new
Abstract: Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later t...
By Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie
arXiv:2609.39350v1 Announce Type: cross
Abstract: As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training para...
By Mengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li, Jing Yang, Tong Yang, Quanlu Zhang, Yu Wang
arXiv:2507.12133v2 Announce Type: replace
Abstract: Device recognition is vital for security in wireless communication systems, particularly for applications like access control. Radio Frequency Fing...
By Hanwen Liu, Yuhe Huang, Yifeng Gong, Yanjie Zhai, Jiaxuan Lu
arXiv:2602.03537v2 Announce Type: replace
Abstract: Matryoshka Quantization (MatQuant), Any-Precision-LLM (AP) and AnyBCQ (AB) are recent quantization approaches showing that a single integer-quantiz...
By Maximilian Kleinegger, Elvir Crn\v{c}evi\'c, Dan Alistarh
arXiv:2604.08454v2 Announce Type: replace
Abstract: Large language models are increasingly deployed in high-stakes domains, where confident yet incorrect inferences may cause severe real-world harm,...
By Haokai Ma, Lee Yan Zhen, Gang Yang, Yunxiang Chen, Yunshan Ma, Tat-Seng Chua, Ee-Chien Chang
arXiv:2605.12904v2 Announce Type: replace
Abstract: Tabular foundation models (TFMs) have emerged as a powerful paradigm for in-context learning on structured data, enabling direct prediction on new...
By Yilong Chen, Xueying Ding, Leman Akoglu
arXiv:2608.13888v2 Announce Type: replace
Abstract: Fashion Outfit Composition (FOC) requires sequentially assembling fashion items into a stylistically cohesive ensemble. Existing works struggle to...
By Kaicheng Pang, Xingxing Zou, Ruohan Xu, Waikeung Wong
arXiv:2609.19011v2 Announce Type: replace
Abstract: Knowledge distillation can copy a deployed model by training a student on its logits or features. The student inherits a watermark only through the...
By Redwanul Karim, Tobias Feigl, Christopher Mutschler, Felix Ott
arXiv:2609.32503v2 Announce Type: replace
Abstract: Kolmogorov-Arnold Networks (KANs) are motivated in part by interpretability: their learned edge functions can be inspected, pruned, and reduced to...
By Ami Tavory, Meir Feder
arXiv:2607.04171v4 Announce Type: replace-cross
Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...
By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv:2607.07494v2 Announce Type: replace-cross
Abstract: Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision fo...
By Jieying Wang, Zizhong Wang, Fangru Linghu, Shuyuan Fan, Jiajia Li, Zhao Zhang
arXiv:2610.00204v1 Announce Type: new
Abstract: Visual-token compression for vision--language models is posed almost entirely as a selection problem: decide which tokens to keep and discard the rest....
By Hongbo Zhang, Zihao Yang, Liuyang Song, Daqian Yang, Haoyang Yao, Yan Wen, Zhengtao Yao
arXiv:2610.00757v1 Announce Type: new
Abstract: Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may re...
By Haowen Guan, Shengzhi Li, Shichao Pei
arXiv:2610.00930v1 Announce Type: new
Abstract: Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating...
By Mingrun Jiang, Yuejia Liu, Zishan Shao, Ting Jiang, Qinsi Wang, Hancheng Ye, Yixiao Wang, Rui-Feng Wang, Kangning Cui, Yixuan Chen, Fan Yang, Xiang Cheng, Hai Li, Yiran Chen
arXiv:2610.01098v1 Announce Type: new
Abstract: Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D recons...
By Hanyuan Xiao, Gonglin Chen, Haolin Xiong, Wenbin Teng, Haiwei Chen, Yajie Zhao
arXiv:2610.00623v1 Announce Type: new
Abstract: Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where li...
By Wenhan Yang, Anirudh Rao, Ashwin Chandra
arXiv:2610.01210v1 Announce Type: new
Abstract: Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, wh...
By Hongming Fu, Jingcheng Shi, Wenjia Wang, Binhua Zuo, Bo Zhao
arXiv:2610.01352v1 Announce Type: new
Abstract: Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven da...
By Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu