Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,725 stories · RSS feed

arXiv Machine Learning
3d ago

Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA

The paper presents a system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media. It tackles three tasks—risk-level classification, evidence phrase extraction, and multi-label factor identification—using Qwen2.5-Instruct models adapted with quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. The final system achieved a composite score of 0.7738, with 0.8089 on Task 1 and 0.6919 on Task 2, demonstrating that task‑specific training and tailored aggregation improve performance across the three tasks.

By Xuan Zhong Feng, Geoffrey Martin, Hexin Dong, Yifan Peng
arXiv Machine Learning
4d ago

JARQ: Joint Alternating Refinement for Quantization

JARQ is a plug‑in refinement method for group‑wise post‑training quantizers that alternates a joint least‑squares fit of all group scales with bounded Babai proposals to move many codes together on the current grid. It solves a bilinear box‑constrained mixed‑integer least‑squares problem without backpropagation, preserves the host model’s bit width, groups, zero points, and inference cost, and improves accuracy across Llama‑2, Llama‑3, and Qwen models when paired with RTN, GPTQ, OmniQuant, and AWQ hosts. In practice, JARQ reduces perplexity in 90 of 96 comparisons, cuts three‑bit RTN perplexity by up to 36%, raises mean multiple‑choice accuracy in 23 of 24 configurations, and enhances QEP, QuaRot, and OJBKQ outputs, all within under a minute per 7B block.

By Xinyu Wang, Sicheng Lyu, Xiao-Wen Chang
arXiv Machine Learning
4d ago

Accelerated Algorithm for Sparse Regularized Partial Optimal Transport

The paper introduces an accelerated algorithm for sparse regularized partial optimal transport (POT). It reformulates the problem using penalty-based techniques that allow efficient gradient updates while maintaining the original structure, and supports a wide range of regularizers that encourage structured sparsity. Experiments on color transfer, domain adaptation, and point cloud registration show the method outperforms existing baselines in transport cost, sparsity, and convergence speed.

By Khoa Nguyen, Dung T. Nguyen, Thong Huynh, Hoang-Hiep Nguyen-Mau, Anh Nguyen, Minh Ngoc Dinh, Juho Kannala
arXiv Machine Learning
4d ago

JET: Justification Evaluation in Transformer

JET (Justification Evaluation in Transformer) leverages pretrained language and vision‑language models to choose among a limited set of answers without extra training. It directly evaluates candidate likelihoods, reuses computation across candidates, and runs experiments on desktop CPUs and consumer GPUs to measure decision accuracy and execution cost. Results show high accuracy on the MMLU test set, significant speedups from prefix reuse and cache management, and a 30.8% reduction in process time through input preparation optimizations, all while maintaining unchanged outputs.

By Shenghao Ding
arXiv Machine Learning
4d ago

Let CSP Be Your ANCHOR: Adaptive Crystal Search over Frozen Structure Priors

The paper proposes ANCHOR, a composition policy that decouples crystal structure prediction (CSP) from de novo crystal generation (DNG) by using a frozen CSP model as a fixed evaluator and training the policy with multi‑objective rewards, including adaptive novelty. ANCHOR significantly improves metrics such as MSUN and SUN compared to traditional DNG approaches, and demonstrates transferability across different CSP backbones and distillation into Crystalite‑CSP. The study shows that fine‑tuning DNG models directly on these rewards yields limited gains, whereas the ANCHOR framework better leverages the CSP prior for more efficient crystal discovery.

By Emma Lei Hovmand, Jonas Elsborg, Melih Kandemir, Arghya Bhowmik
arXiv Computer Vision
4d ago

Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

The paper investigates whether the sparsity of Mixture-of-Experts (MoE) models leads to intrinsic semantic organization across modalities and domains. It shows that experts naturally specialize semantically even without explicit modular training. The authors propose ExpertLens, a data‑free method that decodes router weights to identify domain‑specialized experts, enabling selective fine‑tuning that matches or exceeds full fine‑tuning while updating only 21.7–47.0% of parameters and achieving a 4.0× speedup, outperforming LoRA in both performance and efficiency.

By Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari
arXiv AI
4d ago

Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents

Mem++ is a non‑destructive memory framework for large language model agents that records every document in full, along with its date and author, instead of compressing it at write time. At read time, it retrieves only documents up to the requested time and fuses lexical and semantic rankings, leaving the choice of which version to use to the answering model. Evaluations on OrgMemBench show Mem++ outperforms the strongest baseline by 8.0 to 13.1 points and achieves the best overall score with gpt‑4.1‑mini, while also ranking highly on other benchmarks such as LoCoMo and LongMemEval‑S.

By Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed
arXiv AI
4d ago

Distilling LLM Reasoning into Graph of Concept Predictors

The paper introduces Graph of Concept Predictors (GCP), a reasoning-aware active distillation framework that captures a large language model’s intermediate reasoning as a directed acyclic graph of concepts and mirrors this structure in a smaller student model. GCP improves sample efficiency by using a graph-aware acquisition strategy that weighs concept uncertainty, gradient diversity, and node centrality, and enhances training stability through targeted sub‑module retraining that updates only the most influential concept predictors. Experiments on eight NLP classification benchmarks show that GCP achieves better performance under limited annotation budgets while providing more interpretable and controllable training dynamics.

By Ziyang Yu, Liang Zhao
arXiv Machine Learning
4d ago

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling

DynaFlow is a framework that enables transparent and flexible intra‑device parallelism by separating logical model definitions from physical execution schedules. It offers a flexible frontend with annotations for graph partitioning and a programmable interface for custom parallelism strategies, while its efficient backend handles asynchronous control/data‑flow, custom memory management, and remains compatible with CUDA Graphs and TorchInductor. The authors demonstrate that DynaFlow can integrate representative parallelism strategies into six state‑of‑the‑art ML systems with minimal code changes, achieving up to a 1.29× throughput improvement.

By Yi Pan, Yile Gu, Jinbin Luo, Yibo Wu, Ziren Wang, Hongtao Zhang, Ziyi Xu, Shengkai Lin, Baris Kasikci, Stephanie Wang
arXiv Machine Learning
4d ago

Reinforcement Learning-Guided Graph Transformations for SpTRSV Optimization

The paper introduces a reinforcement learning-guided framework for graph transformations aimed at optimizing sparse triangular solve (SpTRSV) kernels. By treating graph transformation as a sequential decision-making task, the RL agent learns matrix-dependent policies that reduce the number of levels by 23% and the coefficient of variation of level costs by 29%, while modifying less than 1% of matrix rows. Experimental results on real-world sparse matrices show significant performance improvements and demonstrate that the learned policies can transfer to unseen matrices via curriculum learning and fine-tuning.

By Buse Y{\i}lmaz
arXiv AI
4d ago

The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization

The paper investigates how the choice of reconstruction loss, specifically mean squared error (MSE), affects optimization in sequential post‑training quantization (PTQ) of large language models (LLMs). It identifies an Optimization Imbalance where MSE causes reconstruction loss magnitudes—and thus gradient magnitudes—to vary dramatically across quantization stages, leading to uneven parameter updates. The authors propose that root mean squared error (RMSE) variants, which implicitly normalize gradients, can decouple optimization strength from loss scale and outperform MSE as a drop‑in replacement.

By Chao Li, Shigeng Wang, Anbang Yao
arXiv AI
4d ago

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

DMAD (Distribution Matching as Adversarial Distillation) reinterprets distribution matching as a classification task, enabling a few‑step student model to learn log‑density ratios directly via two discriminator heads on a shared backbone. This eliminates the need for an auxiliary diffusion model, reducing memory and computation overhead. Experiments show DMAD achieves state‑of‑the‑art Fréchet Inception Distance scores on ImageNet‑64x64, COCO‑10K, and VBench, and outperforms competing few‑step methods in joint audio‑video generation on MiniMax‑H3‑33B.

By Zhengming Yu, Junkun Yuan, Haotian Yang, Gordon Guocheng Qian, Yizhi Wang, Angtian Wang, Yiding Yang, Bo Liu, Xin Li, Wenping Wang, Chongyang Ma
arXiv AI
4d ago

REALM: Retrospective Encoder Alignment for LFP Modeling

REALM is a retrospective knowledge distillation framework that enables causal decoding of behavior from local field potentials (LFPs). It trains a bidirectional Mamba‑2 teacher on multi‑session data using continuous masked autoencoding, then distills its representations into a compact causal student model. The resulting LFP‑only decoder achieves the highest mean accuracy among compared methods, surpassing state‑of‑the‑art baselines while using fewer parameters and less pretraining time.

By Peicheng Wu, Zhenyu Bu, Runze Ma, Lin Du
arXiv AI
4d ago

TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference

TopK-Guided is a training‑free method that improves activation sparsity for large language model inference by combining token‑level sparsity adaptation with block‑level budget allocation that accounts for block sensitivity. It addresses limitations of existing methods like TEAL, which adapts sparsity per token but lacks tight control, and WINA, which enforces a fixed sparsity across all tokens and blocks. Experiments on Llama‑2 and Llama‑3 show that TopK‑Guided consistently yields better perplexity and downstream accuracy while maintaining similar compute costs to WINA, especially at high sparsity levels.

By Mukund Agarwalla, Chih-Jen Lin
arXiv AI
4d ago

Q-Learning for Reachability in MEC-Free MDPs

The paper introduces Quasar, a model‑free Q‑learning algorithm that guarantees asymptotic convergence for reachability objectives in Markov Decision Processes that are free of non‑terminal maximal end components (MECs). Unlike prior model‑based methods, Quasar does not estimate transition probabilities, reducing memory usage from O(|S|²|A|) to O(|S||A|). Experiments on the Quantitative Verification Benchmark Set show that Quasar converges to optimal policies with far fewer samples than existing state‑of‑the‑art model‑based approaches.

By Lu-Chin Chang, Suguman Bansal
arXiv AI
4d ago

LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation

LEGO-OPD introduces a factorized teacher composition for multimodal on‑policy distillation, combining a Language Expert and a Grounding Expert into a single teacher distribution. By treating the language expert as a prior over tokens and the grounding expert as a visual likelihood that updates this prior, the method decouples language reasoning from visual grounding. Adaptive calibration further adjusts the influence of visual evidence at each decoding prefix, preventing over‑ or under‑supervision. Experiments with Qwen3 models demonstrate that LEGO‑OPD outperforms both single‑ and multi‑teacher baselines on multimodal and text‑only reasoning tasks, improving visual perception while preserving language reasoning.

By Jaeyun Shin, Hangeol Chang, Jong Chul Ye
arXiv AI
4d ago

Manifold-Constrained Initial Noise Optimization for Efficient Generative Model Alignment

The paper introduces ZeNOVA, a gradient‑free method for aligning initial noise in generative models. It uses annealed soft‑value guidance, manifold‑constrained hyperspherical Langevin dynamics, and Metropolis‑Hastings jumps to address instability in black‑box reward settings. Experiments on image and video models show ZeNOVA outperforms existing zeroth‑order baselines by more stably optimizing noise toward higher rewards.

By Jinho Chang, Jong Chul Ye
arXiv AI
4d ago

Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations

This study introduces the first controlled benchmark of generative models for weather data assimilation using real station observations from 11,849 NOAA MADIS stations across the U.S. It evaluates key design choices—diffusion vs. flow matching, pixel vs. latent-space formulations, and inference-time conditioning strategies—against a classical 3D-Var baseline. The benchmark finds that learned generative priors and full-gradient guidance improve RMSE over ERA5, while other design variations offer minimal benefit, especially under sparse observation conditions.

By Ruizhe Huang, Qidong Yang, Jonathan Giezendanner, Sherrie Wang
arXiv AI
4d ago

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

The paper introduces Kinematic MeanFlow (K-MF), a one‑step action generation policy for Robotic Foundation Models that addresses instability in the MeanFlow framework. By decoupling the time derivative into two sub‑interval terms, K-MF captures early and late denoising dynamics separately, reducing error amplification. Experiments show K-MF achieves faster inference—reducing action‑head latency by 67.5%–74.4% and overall end‑to‑end latency by 30.3%–54.9%—while outperforming multi‑step flow matching on various tasks.

By Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao
arXiv AI
4d ago

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

MWOP (Modality-aware Width-wise Operation Pruning) is a method that independently prunes visual‑to‑visual, text‑to‑visual, and text‑to‑text attention paths within each layer of multimodal large language models, and separately selects feed‑forward network channels for visual and textual inputs. It uses a first‑order Taylor criterion to guide pruning, re‑evaluates FFN importance after attention pruning, and applies LoRA‑based recovery training. The approach is paired with path‑sparse Triton attention kernels and compact visual‑side FFN execution to achieve practical acceleration, preserving token sequences while reducing computation. "whyItMatters":"MWOP achieves a 1.6× prefill speedup on LLaVA‑OneVision‑7B while retaining 99.7% performance, and further boosts token‑compression methods to 2.9× and 2.7× speedups, demonstrating its effectiveness across architectures."

By Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen