Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv Computer Vision
Sep 30

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

arXiv:2604.08995v3 Announce Type: replace Abstract: With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, exi...

By Zile Wang, Zexiang Liu, Jiaxing Li, Kaichen Huang, Baixin Xu, Fei Kang, Mengyin An, Peiyu Wang, Biao Jiang, Yichen Wei, Yidan Xietian, Jiangbo Pei, Liang Hu, Boyi Jiang, Hua Xue, Zidong Wang, Haofeng Sun, Wei Li, Wanli Ouyang, Xianglong He, Yang Liu, Yangguang Li, Yahui Zhou
arXiv AI
Sep 30

Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation

The paper investigates how the quantity, source, and selection of prompts influence transfer in on‑policy distillation (OPD) between teacher and student models. It shows that a small set of well‑chosen prompts can achieve performance comparable to large prompt pools, but the effectiveness of prompts depends on the specific teacher‑student pair and target task. The study also finds that prompt utility is relational rather than intrinsic, and that targeted prompt selection does not consistently outperform random sampling.

By Jiaxuan Wang, Jiafei Lyu, Yuchen Cai, Siye Wu, Pengyuan Wang, Jiashun Liu, Xiang Cheng, Kai Yang, Yangkun Chen, Saiyong Yang, Lan-Zhe Guo
arXiv Computer Vision
Sep 30

SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

arXiv:2609.37969v1 Announce Type: new Abstract: High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first genera...

By Haozhe Liu, Tian Ye, Shuchen Xue, Yitong Li, Junsong Chen, Haopeng Li, Jincheng Yu, Duomin Wang, Ruihua Zhang, Lei Zhu, Song Han, Enze Xie