Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,725 stories · RSS feed

arXiv AI
4d ago

Reshape and Recur: Improving SSMs with Input Reshaping and Depth Recurrence

The paper proposes two extensions to State Space Models (SSMs) to reduce memory usage and improve performance. First, it introduces depth recurrence, allowing a looped SSM with fewer parameters to match the performance of a larger, non-recurrent model. Second, it advocates using a fixed time granularity across tasks by reshaping input sequences, which enhances how information is presented to the model. Both techniques consistently benefit four representative SSM architectures (LRU, S5, LinOSS, LrcSSM).

By M\'onika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu
arXiv AI
5d ago

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

The paper proposes a new matrix multiplication approach called SFC-CA GEMM that uses space‑filling curves to partition work in a platform‑ and shape‑oblivious way, achieving communication‑avoiding properties. It demonstrates provable asymptotic communication optimality for both square and rectangular matrices and outperforms vendor libraries on multiple x86 and Arm platforms, with speedups up to 5.5× for specific shapes and 1.8× in weighted harmonic mean throughput. The method is applied to real‑world tasks, improving large‑language‑model inference by up to 1.85× and distributed‑memory GEMM by up to 2.3× over state‑of‑the‑art frameworks.

By Evangelos Georganas, Alexander Heinecke, Pradeep Dubey
arXiv Computer Vision
5d ago

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

The paper introduces a multidimensional observer model that represents images as distributions in a latent perceptual space and models human image quality judgment as comparisons of noisy samples. By aligning the model with neural representations in the primate ventral stream and fitting it to large-scale behavioral data, the authors demonstrate that the perceptual space required for human quality assessment is extremely low-dimensional relative to the image space. The study reveals that the structure of this perceptual space differs between low-level and high-level quality judgments, indicating that humans construct task-dependent perceptual spaces during visual decision making.

By Sheng Zhao, Weikai Lin, Yuhao Zhu
arXiv AI
5d ago

Derandomizing Dense Binary Hypervector Codebooks for Quantized Scalars

The paper introduces a transition-based derandomization framework for dense binary hypervector codebooks used in hyperdimensional computing. It targets two similarity families—exponential and linear decay with scalar separation—and separates the similarity law, derandomization variant, and generator construction. The authors formalize variants that constrain initial Hamming weight, update-count variability, and update balance, deriving exact finite-dimensional expressions for bias, variance, and RMS error, and validate the theory with simulations to guide practical codebook design.

By Dmitri Rachkovskij, Evgeny Osipov, Olexander Volkov, Denis Kleyko, Vaclav Snasel
arXiv Computer Vision
5d ago

DiDA: Video Object Segmentation with Distillation Learning of Deformable Attention

DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.

By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
arXiv Computer Vision
5d ago

Contrastive On-Policy Distillation

Contrastive On-Policy Distillation (COPD) is a framework that improves on-policy distillation by using a frozen teacher to evaluate student states under two contrasting prompts—one encouraging low reasoning effort and one encouraging high effort. The difference in log‑probabilities between these prompts provides a token‑level advantage signal that guides the student toward more concise and efficient reasoning strategies. Experiments on nine multimodal benchmarks show that COPD reduces reasoning length while maintaining task performance, and the contrastive approach can also be applied to on‑policy self‑distillation, allowing a model to compress its own reasoning without an external teacher.

By Jiacheng Ruan, Jun Tang, Wenzhen Yuan, Ting Liu, Shuai Bai, Dayiheng Liu, Zhibo Yang, Yuzhuo Fu
arXiv AI
5d ago

A 3GPP-Compliant Benchmark Dataset for RIS-Aided Beyond 5G Networks

The paper presents a large‑scale, 3GPP TR 38.901‑compliant dataset for RIS‑aided millimeter‑wave B5G networks, covering 20 deployment variants with diverse user densities, fading, and blockage conditions. Each sample includes oracle RIS phase configurations from a brute‑force search, along with full CSI, per‑link channel decomposition, optimal phase matrices, and CQI labels, enabling a wide range of machine‑learning tasks. The authors also introduce a novel CSI‑to‑CQI mapping as a benchmark for scalable link‑quality prediction and evaluate it against state‑of‑the‑art models under various conditions.

By Pujitha Mamillapalli, Pankaj Singh Rathour, Abhinav Kumar
arXiv AI
5d ago

Emergent Multi-View Geometry Through Self-Distillation

The paper introduces Poincar3, a self‑supervised method that learns multi‑view representations through self‑distillation rather than RGB reconstruction. By combining masked patch and image‑level distillation with a teacher that sees additional views, it trains from scratch without explicit 3D supervision. Poincar3 surpasses prior single‑ and multi‑view self‑supervised methods on tasks such as correspondence estimation, camera pose estimation, and 3D reconstruction, and its features encode camera motion more accurately thanks to a lightweight Poincaré adapter.

By David Nordstr\"om, Thibaut Loiseau, Vincent Lepetit, Michael Felsberg, Guillaume Bourmaud, Fredrik Kahl