Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,914 stories · RSS feed

arXiv AI
Sep 15

OCT-FedSIR: Toward Trustworthy Federated Ophthalmic Learning under Annotation Noise

OCT-FedSIR is a reliability‑aware spectral framework designed for federated learning of OCT image classification in the presence of client‑dependent annotation noise and heterogeneous data distributions. It integrates class‑balanced spectral estimation, logit adjustment, complementary spectral descriptors, selective spectral relabeling, and noise‑aware federated optimization. Across 117 experimental conditions on three datasets, OCT‑FedSIR achieved a mean accuracy of 86.73%, outperforming RoFL (79.94%) and FedCorr (78.75%) and successfully identifying and correcting corrupted annotations with high precision.

By Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Rashadul H. Badhon, Behafarin Emam, Sally S. Y. Ong, Atalie C. Thompson, Theodore Leng, Ahmed Arafa, Jennifer I. Lim, Minhaj Nur Alam
arXiv AI
Sep 15

Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning

The paper introduces Verifier-Gated Multi-Expert On-Policy Distillation (VG‑OPD), a method that assigns teacher supervision at the token level based on each expert’s counterfactual gain on a specific answer criterion. VG‑OPD localizes supervision where experts disagree most with the student and weights it by criterion importance, integrating this into a gated KL advantage for reinforcement learning. Applied to scientific reasoning, VG‑OPD achieves top performance on seven benchmarks for 4B and 8B models, outperforming prior multi‑teacher distillation approaches.

By Xun Xu, Zaixi Zhang
arXiv AI
Sep 15

Data-free On-policy Distillation

arXiv:2609.14193v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes...

By Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Jinqiao Wang
arXiv AI
Sep 15

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

ZGCM-1 is a 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. It uses a core premise that compact models can overcome capacity limits by combining deliberate internal thinking with active external tool use, supported by a 256K context and an end‑to‑end high‑efficiency training recipe that includes interleaved gated sliding‑window and full attention, a stable FP8 Muon optimizer, progressive curriculum scaling, and reformulation of interaction traces into Markov Decision Processes. The model is competitive with much larger frontier models on challenging mathematical reasoning and agentic search tasks, offers a ~4.2× efficiency improvement in pre‑training time‑to‑loss, and its weights, checkpoints, training code, data recipes, and logs are fully open‑source to support community research.

By Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren