arXiv:2605.12904v2 Announce Type: replace
Abstract: Tabular foundation models (TFMs) have emerged as a powerful paradigm for in-context learning on structured data, enabling direct prediction on new...
By Yilong Chen, Xueying Ding, Leman Akoglu
arXiv:2608.13888v2 Announce Type: replace
Abstract: Fashion Outfit Composition (FOC) requires sequentially assembling fashion items into a stylistically cohesive ensemble. Existing works struggle to...
By Kaicheng Pang, Xingxing Zou, Ruohan Xu, Waikeung Wong
arXiv:2609.19011v2 Announce Type: replace
Abstract: Knowledge distillation can copy a deployed model by training a student on its logits or features. The student inherits a watermark only through the...
By Redwanul Karim, Tobias Feigl, Christopher Mutschler, Felix Ott
arXiv:2609.32503v2 Announce Type: replace
Abstract: Kolmogorov-Arnold Networks (KANs) are motivated in part by interpretability: their learned edge functions can be inspected, pruned, and reduced to...
By Ami Tavory, Meir Feder
arXiv:2607.04171v4 Announce Type: replace-cross
Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...
By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv:2607.07494v2 Announce Type: replace-cross
Abstract: Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision fo...
By Jieying Wang, Zizhong Wang, Fangru Linghu, Shuyuan Fan, Jiajia Li, Zhao Zhang
arXiv:2610.00204v1 Announce Type: new
Abstract: Visual-token compression for vision--language models is posed almost entirely as a selection problem: decide which tokens to keep and discard the rest....
By Hongbo Zhang, Zihao Yang, Liuyang Song, Daqian Yang, Haoyang Yao, Yan Wen, Zhengtao Yao
arXiv:2610.00757v1 Announce Type: new
Abstract: Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may re...
By Haowen Guan, Shengzhi Li, Shichao Pei
arXiv:2610.00930v1 Announce Type: new
Abstract: Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating...
By Mingrun Jiang, Yuejia Liu, Zishan Shao, Ting Jiang, Qinsi Wang, Hancheng Ye, Yixiao Wang, Rui-Feng Wang, Kangning Cui, Yixuan Chen, Fan Yang, Xiang Cheng, Hai Li, Yiran Chen
arXiv:2610.01098v1 Announce Type: new
Abstract: Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D recons...
By Hanyuan Xiao, Gonglin Chen, Haolin Xiong, Wenbin Teng, Haiwei Chen, Yajie Zhao
arXiv:2610.00623v1 Announce Type: new
Abstract: Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where li...
By Wenhan Yang, Anirudh Rao, Ashwin Chandra
arXiv:2610.01210v1 Announce Type: new
Abstract: Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, wh...
By Hongming Fu, Jingcheng Shi, Wenjia Wang, Binhua Zuo, Bo Zhao
arXiv:2610.01352v1 Announce Type: new
Abstract: Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven da...
By Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu
arXiv:2602.15396v2 Announce Type: replace
Abstract: Diffusion models often yield highly curved trajectories and noisy score targets due to an uninformative, memoryless forward process that induces in...
By Jeongwoo Shin, Jinhwan Sul, Joonseok Lee, Jaewong Choi, Jaemoo Choi
arXiv:2606.21562v2 Announce Type: replace
Abstract: Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming v...
By Philippe Weinzaepfel, Christian Wolf, Mert B\"ulent Sariyildiz, Guillaume Bono, Gianluca Monaci
arXiv:2512.04705v3 Announce Type: replace-cross
Abstract: The deployment of Early-Exiting Neural Networks (EENNs) on edge accelerators requires optimizing not only the network architecture but also i...
By Alaa Zniber, Arne Symons, Ouassim Karrakchou, Marian Verhelst, Mounir Ghogho
DecomVoxel introduces a guided in‑situ denoising optimization that fuses 3D‑native priors with neural scene reconstruction to improve decompositional scene reconstruction. The method employs an epsilon‑based distillation loss for stable latent refinement and adaptive spatial guidance using occupied and vacant anchors with temporal annealing to reduce hallucinations and spatial drift. Experiments on Replica and ScanNet++ demonstrate that DecomVoxel outperforms state‑of‑the‑art approaches while preserving spatial layout, structural fidelity, and style‑consistent texture, yielding high‑quality textured meshes with clean topology.
By Junfeng Ni, Zirui Zhou, Yixin Chen, Yu Liu, Nan Jiang, Zhifei Yang, Song-Chun Zhu, Siyuan Huang
Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.
By Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero
RISED introduces a framework that uses rubric-based textual feedback to improve training of a single large language model (LLM) agent across multiple interactive environments. By having an LLM judge tag rollouts with a shared rubric vocabulary, the system guides both online data selection and policy supervision, enabling richer cross‑environment relationships and within‑group reward contrast. Experiments show that RISED achieves the highest mean pass rate and ranks first or second in every individual environment, with rubric analysis revealing behavioural changes behind these gains.
By Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu
ShatterQuant is a hardware-software co-designed framework that enables mixed-precision quantization within individual tensors by assigning different bit-widths to blocks of a weight projection. It couples precision granularity with processing element configuration, allowing each precision to determine an effective block height. The framework includes a hardware-aware post-training method based on block-level standard deviation and weight sensitivity, a ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, precision-dependent PE configuration, block rescaling, and integrated softmax and piecewise-linear nonlinearities, and an evaluation showing 1.5 TOPS, 760 GOPS/$mm^2$ area efficiency, and 2.8 TOPS/W energy efficiency on a TSMC 16nm PDK implementation.
By Mikolaj Walczak, Edward Humes, Chao Fang, Marian Verhelst, Tinoosh Mohsenin