Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,914 stories · RSS feed

arXiv Computer Vision
Sep 15

MorphoStyle: Motion Style Transfer with Morphology Control

MorphoStyle is a new framework for shape‑aware motion style transfer that uses a shape‑conditioned FSQ‑VAE. It disentangles style from content through a contrastive style encoder, a text‑guided style‑routing mechanism, and a manifold‑preserving style modulator. Experiments on benchmark datasets show that MorphoStyle outperforms existing baselines in both shape control and motion style transfer.

By Xin Feng, Eleonora D'Arnese, Mohan Sridharan
arXiv AI
Sep 15

MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving

MAPS is a Memory-Aware Predictive Scheduling framework designed for disaggregated large language model (LLM) serving. It uses device-assisted speculative output length prediction and uncertainty-aware calibration to establish safe output-length upper bounds, which inform a hierarchical global-local scheduling strategy that reduces queue buildup and head-of-line blocking. Experiments on real-world workloads and two LLMs demonstrate that MAPS lowers average end-to-end latency by 42.6% and tail latency by up to 84.8% compared to three state-of-the-art systems.

By Tiancheng Zhang, Yulin Chen, Yunfeng Zhao, Shaoyuan Huang, Cheng Zhang, Xiaofei Wang
arXiv AI
Sep 15

To See is Not to Master: Teaching LLMs to Use Private Libraries for Code Generation

The paper introduces PriCoder, a method for teaching large language models (LLMs) to effectively use private library APIs for code generation. PriCoder synthesizes training data by constructing a graph and applying two operators—Progressive Graph Evolution to increase diversity and Multidimensional Graph Pruning to enhance quality. Experiments on three mainstream LLMs demonstrate that PriCoder boosts private‑library code generation by over 20% in pass@1, while leaving general code generation largely unchanged.

By Yitong Zhang, Chengze Li, Ruize Chen, Guowei Yang, Xiaoran Jia, Yijie Ren, Jia Li
arXiv AI
Sep 15

MANE: A Multi-Path Adaptive Network for Edge Onloading of Deep Neural Networks

MANE is a distributed inference framework that uses a multi‑path tail architecture to allow dynamic accuracy–throughput trade‑offs during edge onloading of deep neural networks. It introduces a novel multi‑path model, a three‑stage training scheme with Joint Head Network Distillation loss, and a hysteresis‑based scheduler with an equitable device‑fallback policy. The system achieves over 80% SLO satisfaction and 6pp higher accuracy than on‑device alternatives while supporting up to 40 concurrent devices.

By Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias, Iakovos S. Venieris
arXiv AI
Sep 15

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction

The paper investigates how much learned memory is required to leverage additional data in autoregressive prediction models. It introduces a predictive‑energy spectrum that jointly governs data and memory scaling, proving a minimax law that links the number of prediction blocks and the size of the learned state to this spectrum. The authors demonstrate that optimal bit allocation and masked query‑key attention mechanisms realize this law, and they provide experimental evidence across multiple pretrained‑model scales.

By Chiwun Yang, Xiaoyu Li
arXiv AI
Sep 15

Tensorization is a powerful but underexplored tool for compression and interpretability of neural networks

The paper discusses tensorizing neural networks by reshaping dense weight matrices into higher-order tensors and approximating them with low-rank tensor network decompositions. This approach offers promising model compression and introduces bond indices that create new latent spaces, potentially enhancing interpretability. Despite encouraging empirical results, tensorized neural networks remain underused, and the authors call for more research to address practical scaling and adoption challenges.

By Safa Hamreras, Sukhbinder Singh, Rom\'an Or\'us
arXiv AI
Sep 15

LLMs or Naive Bayes? Old Gems or New Ways

The paper compares Complement Naive Bayes (NB) with zero‑shot and few‑shot large language models (LLMs) across a wide range of model sizes and text classification tasks. NB outperforms LLMs when labeled data is available, achieving comparable accuracy to large LLMs while running thousands of samples per second on a CPU. In zero‑data sentiment settings, LLMs still dominate, but NB remains the best choice for resource‑constrained HPC practitioners, and the authors provide a Kubernetes Helm operator to automate model selection.

By Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank W\"urthwein
arXiv Machine Learning
Sep 15

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

arXiv:2609.15972v1 Announce Type: cross Abstract: As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the p...

By Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang
arXiv AI
Sep 15

An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

The paper introduces a modular correction framework for large language models that uses Activated LoRA adapters and a context-aware routing mechanism to mitigate harmful outputs. By allowing expert adapters to activate mid-sequence without invalidating the KV cache, the system achieves low-latency, targeted correction during generation. Experiments show improved alignment on safety benchmarks while maintaining task performance, presenting a lightweight, scalable approach to safer LLM deployments.

By Roberto Campbell, Momin Abbass, Muneeza Azmat, Michal Ulewicz, Raya Horesh, Kristjan Greenewald, Rog\'erio Abreu de Paula, Nathalie Baracaldo