Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Machine Learning
Sep 24

hyperbolix: Hyperbolic Deep Learning in JAX

hyperbolix is an open‑source library for hyperbolic deep learning in JAX, built on Flax NNX. It provides six manifolds—including Euclidean, Poincaré ball, hyperboloid, κ‑stereographic, mixed‑curvature product, and proper velocity space—through a common interface, and implements a wide range of layer families (linear, convolution, attention, normalization, positional encoding, regression, vector quantization). The library also supplies Riemannian optimizers, wrapped distributions, dimensionality‑reduction techniques, and precision‑tested operations that replace numerically unstable formulas on the hyperboloid, ensuring accurate float32 computations at large distances.

By Timo Klein, Thomas Lang, Yllka Velaj, Sebastian Tschiatschek
arXiv Machine Learning
Sep 24

Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections

The paper investigates minimal‑norm interpolation and λ2‑regularized logistic‑loss minimization for binary classification using univariate two‑layer ReLU networks. It provides exact geometric characterizations of optimal classifiers, showing that unpenalized hidden‑layer biases yield continuous piecewise‑affine functions that tightly follow label switches, while penalized biases produce a unique, sparsest classifier with a single kink per same‑label segment. Adding a free affine skip connection does not change these function‑space solutions but guarantees that every KKT point becomes globally optimal, eliminating suboptimal KKT points that can arise without the skip connection.

By Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c
arXiv Machine Learning
Sep 24

The Computational Value of Sensory-Aligned Receptive Fields Depends on Neuronal Expressivity

The study investigates whether sensory-aligned receptive fields provide computational benefits beyond mere resource efficiency in recurrent networks of Expressive Leaky Memory neurons. Across auditory and event-based visual classification tasks, receptive fields aligned with task-relevant sensory coordinates improve test accuracy compared to budget-matched random fields, but this advantage disappears when coordinates are scrambled or irrelevant. The benefit diminishes as neuronal expressivity increases, and generic synaptic sparsity regularization only partially recovers performance, indicating that structured receptive fields act as a computational prior beyond sparsity alone.

By Agnese Adorante, Aaron Spieler, Anna Levina
arXiv AI
Sep 24

Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching

The paper proposes a method to train efficient multi‑task manipulation policies by distilling knowledge from single‑task Conditional Flow Matching (CFM) experts. Instead of training separate models for each task, the authors transfer the experts’ learned velocity fields into a shared policy, combining this distillation signal with the original CFM objective. Experiments on RLBench demonstrate that this approach improves multi‑task performance while keeping the model size fixed, avoiding the need for larger capacity or performance drops seen with naive concatenated training.

By Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada
arXiv Machine Learning
Sep 24

NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers

The paper introduces Newton-Schulz Attention (NS-Attn.), a parameter‑free transformation that applies a finite Newton‑Schulz polynomial step to the output of each attention head in Vision Transformers. By normalizing each head’s feature‑by‑token matrix with its Frobenius norm, applying the NS step, and restoring the norm, the method aims to reduce spectral concentration and increase effective rank before merging heads. Experiments on ViT and Swin models over CIFAR‑10 and CIFAR‑100 show consistent accuracy gains of 0.25–0.83 percentage points, though with added inference latency.

By Xiaohe Jiang (University of Exeter), Guoqiang Zhang (University of Exeter), Tianjin Huang (University of Exeter), Ronghui Mu (University of Exeter)
arXiv AI
Sep 24

Hunyuan-A13B Technical Report

arXiv:2609.27284v1 Announce Type: new Abstract: We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but ac...

By Tencent Hunyuan Team, Ao Liu, Botong Zhou, Can Xu, Chayse Zhou, ChenChen Zhang, Chengcheng Xu, Chenhao Wang, Decheng Wu, Dengpeng Wu, Dian Jiao, Dong Du, Dong Wang, Feng Zhang, Fengzong Lian, Guanghui Xu, Guanwei Zhang, Hai Wang, Haipeng Luo, Han Hu, Huilin Xu, Jiajia Wu, Jianchen Zhu, Jianfeng Yan, Jiaqi Zhu, Jihong Zhang, Jinbao Xue, Jun Xia, Junqiang Zheng, Kai Liu, Kai Zhang, Kai Zheng, Kejiao Li, Keyao Wang, Lan Jiang, Lixin Liu, Lulu Wu, Mengyuan Huang, Peijie Yu, Peiqi Wang, Qian Wang, Qianbiao Xiang, Qibin Liu, Qingfeng Sun, Richard Guo, Ruobing Xie, Saiyong Yang, Shaohua Chen, Shihui Hu, Shuai Li, Shuaipeng Li, Shuang Chen, Suncong Zheng, Tao Yang, Tian Zhang, Tinghao Yu, Weidong Han, Weijie Liu, Weijin Zhou, Weikang Wang, Wesleye Chen, Xiao Feng, Xiaoqin Ren, Xingwu Sun, Xiong Kuang, Xuemeng Huang, Xun Cao, Yanfeng Chen, Yang Du, Zhen Yang, Yangyu Tao, Yaping Deng, Yi Shen, Yigeng Hong, Yiqi Chen
arXiv Machine Learning
Sep 24

Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following

arXiv:2609.27421v1 Announce Type: new Abstract: Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tun...

By Yanzhao Zheng, Yuanqiang Yu, Tianze Xu, Chao Ma, Zhentao Zhang, Jihuai Zhu, Baohua Dong, Hangcheng Zhu, Ruohui Huang