Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

6,032 stories · RSS feed

arXiv Machine Learning
Sep 10

Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems

The paper reports an empirical scalability study of data‑parallel training for Kolmogorov‑Arnold Networks (KANs) on high‑performance computing systems. Using up to eight NVIDIA A100 GPUs across four nodes on the FinisTerrae III supercomputer, the authors evaluate strong and weak scaling, communication overhead, and model‑size scaling, finding a 74.7% parallel efficiency and a 5.97× speedup at eight GPUs. They observe non‑monotonic communication costs driven by All‑Reduce choices and inter‑node latency, and note that while the parameter‑to‑memory ratio improves with larger models, training time scales less favorably, leading to guidelines for GPU topology and model‑size selection.

By Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso
arXiv Machine Learning
Sep 10

Large-Scale Pretraining for Improving Deep Learning-Based Geometric Distortion Correction of Diffusion-Weighted Imaging

The paper explores large‑scale pretraining to enhance deep learning‑based geometric distortion correction for diffusion‑weighted imaging (DWI). By framing the task as image reconstruction, the authors compare a non‑pretrained baseline with self‑supervised and generative pretrained models, finding that the cWDM model yields the best quantitative and qualitative results. When applied to low‑resource, high‑throughput settings in a low‑ and middle‑income country, the pretrained models faced transferability issues, but aligning images to a common standard space improved predictions, indicating that harmonized preprocessing can aid cross‑domain deployment.

By Saroj Khanal, Yashawant Kumar Yadav, Kritam Bhattarai, Jeevan Neupane, Shristi Subedi, Saship Gwachha, Manish Kumar Tiwari, Dong Zhang, Confidence Raymond, Aondona Moses Iorumbur, Udunna Anazodo, Surendra Maharjan, Bishesh Khanal, Mahesh Shakya, Pralhad Kumar Shrestha
arXiv Machine Learning
Sep 10

PatchFormer: A Patch-Based Time Series Foundation Model with Hierarchical Masked Reconstruction and Cross-Domain Transfer Learning for Zero-Shot Multi-Horizon Forecasting

arXiv:2601.20845v2 Announce Type: replace Abstract: Time series forecasting is a fundamental problem with applications in climate, energy, healthcare, and finance. Many existing approaches require do...

By Olaf Yunus Laitinen Imanov, Derya Umut Kulali, Taner Yilmaz
arXiv Machine Learning
Sep 10

Miles v0.1: Production-Level Post-Training

arXiv:2609.08368v1 Announce Type: new Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each st...

By RadixArk, :, Tom Chen, Mao Cheng, Shi Dong, Kangrui Du, Yanbin Jiang, Jiajun Li, Yiming Li, Tao Lin, Yusheng Su, Andy Ye, Yueming Yuan, Zhichen Zeng