Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,725 stories · RSS feed

arXiv AI
4d ago

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Selection-Based Structured Reasoning (SSR) is a framework that replaces free-form reasoning with the selection of pre-specified natural‑language reasoning candidates, allowing multimodal agents to choose from a set of reusable high‑level reasoning traces. By scoring these candidates in parallel using a shared context KV cache, SSR eliminates the need for an auxiliary task head and reduces inference cost. Experiments on seven multimodal search benchmarks with 2B and 4B models show that SSR maintains competitive success rates while cutting per‑turn reasoning latency by over 90% and overall inference latency by 28–54%.

By Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
arXiv AI
4d ago

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

The paper introduces GALA, a distillation technique that replaces costly neural decoding in 3D Gaussian avatars with a shallow MLP predicting blendshape coefficients, enabling real‑time animation. By constructing a basis via block‑local PCA under a rendering‑aware metric, GALA achieves high fidelity while reducing memory usage. Experiments on three avatar models show up to three orders of magnitude lower CPU cost and frame rates up to 60fps on mobile devices.

By Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
arXiv AI
4d ago

Masked Self-Distillation: Internalizing the Chain-of-Thought in Language Models

The paper introduces masked self‑distillation, a post‑training framework that trains a language model to internalize portions of its own intermediate reasoning traces. By varying the fraction of trace internalized, the authors demonstrate that models can achieve higher inference efficiency and improved task performance on math and graph‑coloring problems. Experiments on Qwen3‑4B and Qwen3‑8B show that the method generalizes well to in‑domain out‑of‑distribution cases without catastrophic forgetting, and that supervised fine‑tuning alone can reduce trace length at the expense of generalization.

By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv AI
4d ago

Pepti-drift: Scalable Safe-Active Peptide Generation Without Inference-Time Guidance

Pepti-drift is a one‑step peptide generation framework that uses a single latent refinement and parallel decoding, avoiding costly inference‑time guidance. It incorporates attraction toward target‑specific binders and repulsion from liability‑associated regions, achieving an 18.37% predicted Safe‑Active yield across 88 held‑out targets. The method delivers competitive gains to multi‑property‑guided baselines while reducing generation cost by 468×, enabling scalable and fair high‑throughput peptide design.

By Takashi Fujiwara, Hikaru Shindo, Kaushalya Madhawa, Jun Jin Choong, Shuan Chen, Yuna Oikawa, Yiming Zhang, Gyubok Lee, Keisuke Ozawa
arXiv Machine Learning
4d ago

Activation-Conditioned Self-Distillation

Activation-Conditioned Self-Distillation (ACSD) is a new on‑policy self‑distillation method that uses a frozen copy of the base model to extract a steering vector by contrasting activations from self‑generated trajectories that reach verified correct answers with all other trajectories. The student learns from next‑token distributions on its own prefixes, without needing reference text or teacher parameter updates, and is used alone at inference. Across five models, ACSD achieves the highest mean accuracy on four mathematical benchmarks, with notable gains on DeepSeek‑R1‑0528‑Qwen3‑8B and LiveCodeBench v6 compared to the OPSD baseline.

By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv Machine Learning
4d ago

SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs

The paper introduces SparLeak, a side‑channel attack that exploits a new GPU micro‑architectural leakage called Sparsity‑Induced Memory Access (SIMA) caused by sparse attention in large language models. By capturing SIMA traces during LLM inference, SparLeak can infer user‑query attributes from prefill‑phase traces and reconstruct autoregressive responses from decoding‑phase traces. Experiments on three LLM architectures, three sparse attention mechanisms, and three privacy‑sensitive datasets show high success rates (90.9% for attribute inference and 87.3% for response reconstruction), underscoring the need to address SIMA leakage in sparse‑attention deployments.

By Fahao Chen, Linkang Du, Jinhao Zhou, Peng Li, Zhou Su
arXiv Machine Learning
4d ago

Visualizing Distribution Coverage in Generative Diffusion Models

The paper introduces pass@k, a metric that evaluates how well generative diffusion models cover their output distribution by measuring the probability that at least one of k independent samples meets a quality criterion. Using this metric, the authors show that while classifier‑free guidance improves single‑draw quality, its advantage diminishes or reverses as k increases, revealing a trade‑off between quality and distribution coverage. They further demonstrate that the choice of training objective in diffusion distillation—whether distribution‑matching or consistency/trajectory‑based—determines whether a few‑step model retains its teacher’s coverage or sacrifices it for higher single‑draw performance, a phenomenon that also appears in few‑step causal video generation.

By Yifei Wang, Xiaoyu Wu, Tsu-Jui Fu, Chen Chen, Liang-Chieh Chen, Zhe Gan, Chen Wei
arXiv Machine Learning
4d ago

Smaller Models, Better Rejects: Preference Distillation Scaling

The paper challenges two common assumptions in preference distillation: that self-generated failures are the best negatives and that rejects must come from large models. Experiments show that smaller frozen models can generate high‑quality rejects with less compute, improving student performance on code generation and math reasoning. The authors provide a theoretical bound on Direct Preference Optimization, identify three practical interventions—mixing rejects, shuffling tokens, and selecting low‑likelihood candidates—that further enhance reject utility, and argue that task structure, not just reference policy coupling, drives effectiveness.

By Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao
arXiv Machine Learning
4d ago

ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

The paper introduces ID Balancing, an Integral‑Derivative controller that improves expert load balance in extremely sparse Mixture‑of‑Experts (MoE) models. By scaling its integral term with load error and activating the derivative term only when imbalance worsens, ID Balancing achieves over 50% reduction in worst‑case backbone MaxVio and 12% reduction in training‑average backbone MinVio compared to leading baselines, while preserving language‑modeling performance across various routing settings. The method’s benefits grow with increased sparsity, making it a promising approach for scaling larger MoE models.

By Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu
arXiv Machine Learning
4d ago

Low-Discrepancy Dither for Quantized Recurrent State Caches

The paper investigates rounding strategies for low‑precision recurrent state caches in Mamba‑style and hybrid language models. It shows that a deterministic golden‑ratio Weyl dither consistently yields quantized models closer to full precision than stochastic rounding, across various models, storage formats, and long decoding horizons, without extra cost. In contrast, round‑to‑nearest can appear effective in short tests but degrades over long generations due to error accumulation.

By Snigdha Chandan Khilar
arXiv Machine Learning
4d ago

MANET-GNN: Learned Decentralized Optimization of Power Allocation in Multi-Channel MANETs

The paper introduces MANET-GNN, a message‑passing graph neural network designed to perform decentralized power allocation in dynamic, multi‑hop, multi‑channel mobile ad hoc networks (MANETs). It formulates a constrained throughput maximization problem that includes various traffic patterns such as unicast, multicast, and convergecast, and uses this as an unsupervised training objective. MANET‑GNN operates with only local, possibly noisy channel state information and a limited number of neighbor exchanges, achieving performance comparable to centralized solutions while scaling across different network topologies and sizes.

By Tomer Alter, Nir Shlezinger, Michael Segal
arXiv Machine Learning
4d ago

PMosFM: Preconditioned Manifold Matching for One-Step Physics-Constrained Generation

PMosFM introduces a preconditioned manifold matching framework that enables one‑step physics‑constrained generation by encoding constraints in a manifold decoder. The method learns transport in intrinsic coordinates, eliminating the need for residual losses or trajectory unrolling, and employs a geometric preconditioner and covariance transform to improve conditioning. Experiments demonstrate reduced training and sampling time compared to multi‑step baselines while maintaining comparable physical and distributional fidelity.

By Zhangyong Liang, Haibin Ling
arXiv Machine Learning
4d ago

Compression Footprints as Security Signals for Model-Poisoning Defense in Federated Learning

The paper proposes using lossy compression as a security signal in federated learning by defining a "compression footprint"—a low‑dimensional set of statistics derived from the compressor’s output. It introduces CRAFT, a server‑side aggregation method that leverages these footprints to distinguish honest from malicious updates without extra client metadata or knowledge of attacker numbers. Experiments on six attacks and datasets show CRAFT often outperforms existing robust aggregation baselines, demonstrating that compression can simultaneously reduce communication and enhance security.

By Sachi Shome, William Eiers
arXiv Machine Learning
4d ago

Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse

The paper presents a study on securing large language model (LLM) inference services against adversarial use, such as jailbreaking, denial‑of‑service, and distillation attacks. It introduces a structural causal model to generate a realistic, labeled dataset of user sessions, including coordinated multi‑account campaigns and varying label observability. Using this dataset, the authors train a gradient‑boosted detector that achieves near‑perfect binary classification on oracle labels but shows lower performance on operational labels, and they demonstrate a simple thresholded engine that improves attack‑type attribution accuracy.

By Keifer Lee
arXiv Machine Learning
4d ago

Model-to-Data Distillation for Graph Neural Networks

The paper introduces Model-to-Data (M2D) distillation, a method that transfers properties learned by a complex graph neural network (GNN) teacher into the graph data itself. By jointly learning augmented node features and graph structure, M2D encodes the teacher’s behavior, allowing simpler GNNs to recover high predictive performance, fairness, and robustness. The resulting graph can be reused with various downstream models, enabling them to approximate sophisticated teachers such as fairness-aware GNNs, Graph Attention Networks, and Graph Transformers.

By Debolina Halder Lina, Arlei Silva
arXiv Machine Learning
4d ago

Surprisingly High Redundancy in Electronic Structure Data Across Materials Explained by Low Intrinsic Dimensionality

The study shows that electronic structure datasets used for machine learning contain significant redundancy due to low intrinsic dimensionality. Random pruning and a coverage-based strategy can reduce dataset size by up to two orders of magnitude while preserving chemical accuracy and generalizability, cutting training time by a factor of three or more. The essential information lies on a low‑dimensional, non‑linear manifold, suggesting that large datasets may be largely overlapping and that minimal, representative datasets could suffice for accurate predictions.

By Sazzad Hossain, Ponkrshnan Thiagarajan, Shashank Pathrudkar, Stephanie Taylor, Abhijeet S. Gangan, Amartya S. Banerjee, Susanta Ghosh
arXiv Computer Vision
4d ago

PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

PAGER is a label‑free adaptation method that aligns partial, viewpoint‑dependent 3D observations with a frozen global semantic space. It uses matched‑point feature alignment and relational supervision to anchor partial features to their global counterparts while preserving similarity structure, all without altering the pretrained encoder or global probe. Experiments show PAGER outperforms label‑supervised PEFT on Sonata and Concerto, and achieves superior zero‑shot transfer from ScanNet to ScanNet++ compared to fully fine‑tuned Sonata.

By Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni
arXiv Computer Vision
4d ago

Towards Fast and Disentangled Counterfactuals for Visual Foundation Models

The paper introduces Disentangled Diffusion Autoencoders (DiDAE), a method that wraps a frozen foundation model with a conditional diffusion decoder to generate counterfactual edits along disentangled dictionary directions. DiDAE can use supervised or unsupervised dictionaries, requires no gradients, and is up to 2000 times faster than existing approaches. Evaluations on six datasets show that its counterfactuals match or surpass state‑of‑the‑art methods and can repair downstream classifiers via Counterfactual Knowledge Distillation, with the entire workflow available as an open‑source library.

By Sidney Bender, Benedikt Kunz, Ahmed Zeid, Shinichi Nakajima, Klaus-Robert M\"uller, Marco Morik
arXiv Statistics ML
4d ago

Unbiased Top-$k$ Estimation for On-Policy Distillation

The paper introduces Tail‑Corrected Top‑k On‑Policy Distillation (TT‑OPD), a method that improves on existing Top‑k OPD by combining the selected top‑k tokens with a sampled token from the student’s rollout. This hybrid approach recovers the probability mass discarded by limiting to top‑k, yielding an unbiased estimator of the reverse KL divergence gradient while maintaining low computational cost. Experiments show TT‑OPD outperforms other OPD variants in accuracy.

By Linjian Meng, Siyuan Gan, YuHan Li, Xiran Wang, Ziyang Ding, Ditang Gou, Yiming Wu, Zhen Zhao
arXiv AI
4d ago

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation

The paper introduces Hierarchical Hash Retrieval (HHR), a coarse‑to‑fine framework designed to improve hash‑based retrieval for large language models. HHR combines Geometry‑Aware Key Routing (GKR) to redistribute feature magnitudes and prune low‑logit keys, with Learned Hash Projection (LHP) to align Hamming distance with true query‑key relevance for fine‑grained retrieval. Experiments on diverse LLMs and benchmarks show that HHR outperforms existing methods, boosting LongBench scores by 1.10 points and achieving up to 3.30× decoding speedup at 128K context length for Llama‑3.1‑8B‑Instruct.

By Lianjun Liu, Tiantian Zheng, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong