Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Computer Vision
Sep 22

What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization

arXiv:2609.24691v1 Announce Type: new Abstract: Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the late...

By Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa, Julian McGinnis, Jiancheng Yang, Daniel Rueckert, Jiazhen Pan
arXiv Machine Learning
Sep 22

A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting

The paper introduces a hybrid attention model that learns a unified time‑aware patch representation for irregular multivariate time series (IMTS) forecasting. It employs a time‑aware patch encoding to embed variable‑length intra‑patch timestamps, a time bias attention mechanism to adjust for temporal misalignment and asynchronous cross‑channel dependencies, and a hybrid causal mask on a decoder‑only Transformer to balance historical context with autoregressive forecasting. The authors also curate VersaTSA, a 30 B‑observation dataset preserving native sampling sparsity, and demonstrate state‑of‑the‑art zero‑shot performance on three IMTS benchmarks while remaining competitive on regular MTS tasks.

By Zhihao Lin, Li Lin, Qi Zhang, Kaiwen Xia, Shuai Wang, Jialin Qiao
arXiv Machine Learning
Sep 22

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

arXiv:2609.23697v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...

By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang