Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

6,032 stories · RSS feed

arXiv AI
Sep 15

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

The paper introduces Counteraction-Aware Multi-Teacher On-Policy Distillation (CaMOPD), a method designed to recover general capabilities in large language models while preserving domain-specific behavior. CaMOPD tackles two failure modes of standard Multi-Teacher On-Policy Distillation—conflicting recovery and preservation gradients, and weak correction signals—by using decoupled alternating training and selecting samples with large teacher‑student log‑probability gaps. Experiments on role‑play dialogue and medical reasoning QA show that CaMOPD outperforms baselines in general capability recovery while maintaining domain specialization, and gradient coherence analyses confirm more coherent correction signals.

By Tianlei Chen, Jiao Ou, Ziyuan Liu, Ruiming Tang, Jian Liang, Han Li
arXiv Machine Learning
Sep 15

Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication

The paper proposes a privacy‑aligned personalized federated learning method that releases a private client context once and limits repeated adaptation to a fixed coefficient space, thereby reducing dimensionality misalignment. A factorized generator creates an adaptive optimization geometry that reshapes noisy updates, and most of the private‑training benefit is preserved by radial evolution. Variable‑length Gaussian quantization is used for coefficient updates, allowing the quantization error to act as the privacy perturbation and cutting protected uplink communication by a factor of 2.67 on CIFAR‑10 at ε=16 while maintaining comparable future‑client accuracy.

By Yilin Xu, Chun Hei Michael Shiu, Chih Wei Ling, Linqi Song
arXiv Computer Vision
Sep 15

MorphoStyle: Motion Style Transfer with Morphology Control

MorphoStyle is a new framework for shape‑aware motion style transfer that uses a shape‑conditioned FSQ‑VAE. It disentangles style from content through a contrastive style encoder, a text‑guided style‑routing mechanism, and a manifold‑preserving style modulator. Experiments on benchmark datasets show that MorphoStyle outperforms existing baselines in both shape control and motion style transfer.

By Xin Feng, Eleonora D'Arnese, Mohan Sridharan