The paper proposes two extensions to State Space Models (SSMs) to reduce memory usage and improve performance. First, it introduces depth recurrence, allowing a looped SSM with fewer parameters to match the performance of a larger, non-recurrent model. Second, it advocates using a fixed time granularity across tasks by reshaping input sequences, which enhances how information is presented to the model. Both techniques consistently benefit four representative SSM architectures (LRU, S5, LinOSS, LrcSSM).
By M\'onika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu
arXiv:2609.31882v2 Announce Type: replace-cross
Abstract: Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to esti...
By Zhengyi Guo, Jiayuan Sheng, Wenpin Tang, David D. Yao
arXiv:2404.11624v3 Announce Type: replace-cross
Abstract: We introduce Token Space, a categorical framework for AI computations based on explicit structural records. Five theses guide it: object inte...
By Wuming Pan
arXiv:2609.32353v2 Announce Type: replace
Abstract: Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extre...
By Junxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu, Weisheng Dong, Yulun Zhang
Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction,...
The paper proposes a new matrix multiplication approach called SFC-CA GEMM that uses space‑filling curves to partition work in a platform‑ and shape‑oblivious way, achieving communication‑avoiding properties. It demonstrates provable asymptotic communication optimality for both square and rectangular matrices and outperforms vendor libraries on multiple x86 and Arm platforms, with speedups up to 5.5× for specific shapes and 1.8× in weighted harmonic mean throughput. The method is applied to real‑world tasks, improving large‑language‑model inference by up to 1.85× and distributed‑memory GEMM by up to 2.3× over state‑of‑the‑art frameworks.
By Evangelos Georganas, Alexander Heinecke, Pradeep Dubey
The paper introduces a multidimensional observer model that represents images as distributions in a latent perceptual space and models human image quality judgment as comparisons of noisy samples. By aligning the model with neural representations in the primate ventral stream and fitting it to large-scale behavioral data, the authors demonstrate that the perceptual space required for human quality assessment is extremely low-dimensional relative to the image space. The study reveals that the structure of this perceptual space differs between low-level and high-level quality judgments, indicating that humans construct task-dependent perceptual spaces during visual decision making.
By Sheng Zhao, Weikai Lin, Yuhao Zhu
The paper introduces a transition-based derandomization framework for dense binary hypervector codebooks used in hyperdimensional computing. It targets two similarity families—exponential and linear decay with scalar separation—and separates the similarity law, derandomization variant, and generator construction. The authors formalize variants that constrain initial Hamming weight, update-count variability, and update balance, deriving exact finite-dimensional expressions for bias, variance, and RMS error, and validate the theory with simulations to guide practical codebook design.
By Dmitri Rachkovskij, Evgeny Osipov, Olexander Volkov, Denis Kleyko, Vaclav Snasel
DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.
By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
Contrastive On-Policy Distillation (COPD) is a framework that improves on-policy distillation by using a frozen teacher to evaluate student states under two contrasting prompts—one encouraging low reasoning effort and one encouraging high effort. The difference in log‑probabilities between these prompts provides a token‑level advantage signal that guides the student toward more concise and efficient reasoning strategies. Experiments on nine multimodal benchmarks show that COPD reduces reasoning length while maintaining task performance, and the contrastive approach can also be applied to on‑policy self‑distillation, allowing a model to compress its own reasoning without an external teacher.
By Jiacheng Ruan, Jun Tang, Wenzhen Yuan, Ting Liu, Shuai Bai, Dayiheng Liu, Zhibo Yang, Yuzhuo Fu
The paper presents a large‑scale, 3GPP TR 38.901‑compliant dataset for RIS‑aided millimeter‑wave B5G networks, covering 20 deployment variants with diverse user densities, fading, and blockage conditions. Each sample includes oracle RIS phase configurations from a brute‑force search, along with full CSI, per‑link channel decomposition, optimal phase matrices, and CQI labels, enabling a wide range of machine‑learning tasks. The authors also introduce a novel CSI‑to‑CQI mapping as a benchmark for scalable link‑quality prediction and evaluate it against state‑of‑the‑art models under various conditions.
By Pujitha Mamillapalli, Pankaj Singh Rathour, Abhinav Kumar
The paper introduces Poincar3, a self‑supervised method that learns multi‑view representations through self‑distillation rather than RGB reconstruction. By combining masked patch and image‑level distillation with a teacher that sees additional views, it trains from scratch without explicit 3D supervision. Poincar3 surpasses prior single‑ and multi‑view self‑supervised methods on tasks such as correspondence estimation, camera pose estimation, and 3D reconstruction, and its features encode camera motion more accurately thanks to a lightweight Poincaré adapter.
By David Nordstr\"om, Thibaut Loiseau, Vincent Lepetit, Michael Felsberg, Guillaume Bourmaud, Fredrik Kahl
arXiv:2607.13808v2 Announce Type: replace
Abstract: Neural radiance representations in Gaussian Splatting (GS) deliver high-fidelity color detail but impose substantial rendering overhead from networ...
By Neel Kelkar, Simon Niedermayr, Kaloian Petkov, Klaus Engel, R\"udiger Westermann
arXiv:2609.38697v1 Announce Type: cross
Abstract: We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources...
By Matias Parij, Pawan Paudel, Tate Berenbaum, Muthaiah Venkatachalam
arXiv:2603.01623v2 Announce Type: replace
Abstract: Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference spe...
By Jiaqi Han, Juntong Shi, Puheng Li, Haotian Ye, Qiushan Guo, Stefano Ermon
arXiv:2609.39021v1 Announce Type: new
Abstract: Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly...
By Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang
arXiv:2609.39132v1 Announce Type: new
Abstract: We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substanti...
By Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng
arXiv:2609.39222v1 Announce Type: new
Abstract: High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff be...
By Xu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou
arXiv:2609.40333v1 Announce Type: new
Abstract: Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are in...
By Ivan Martinovi\'c, Lukas Knobel, Yuki M. Asano
arXiv:2609.35490v2 Announce Type: replace
Abstract: Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning....
By Mohammad Mahdi, Nedyalko Prisadnikov, Yuqian Fu, Carmelo Scribano, Danda Pani Paudel, Luc Van Gool