Hugging Face Blog

Rocket Money x Hugging Face: Scaling Volatile ML Models in Production​

arXiv Machine Learning
Sep 22

Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets

Linear models with an $L_1$-norm penalty are still the leading approach for high‑dimensional tasks, yet many existing solvers are slow, ineffective, and hard to parallelise, making them unsuitable for large industry‑scale corpora. The paper evaluates several recent state‑of‑the‑art methods and shows that older techniques outperform them in general use. It also demonstrates that a simple baseline—LBFGS applied to a sub‑gradient with minor tweaks—yields strong performance and is easier to support and scale in production.

By Edward Raff, James Holt
arXiv AI
Jun 8

Design Once, Deploy at Scale: Template-Driven ML Development for Large Model Ecosystems

arXiv:2603. 24963v3 Announce Type: replace Abstract: Modern computational advertising platforms typically rely on recommendation systems to predict user responses, such as click-through rates, conversion rates, and other optimization events.

By Jiang Liu, John Martabano Landy, Yao Xuan, Swamy Muddu, Nhat Le, Munaf Sahaf, Luc Kien Hang, Rupinder Khandpour, Kevin De Angeli, Chang Yang, Shouyuan Chen, Shiblee Sadik, Anirudh Agrawal, Djordje Gligorijevic, Jingzheng Qin, Peggy Yao, Alireza Vahdatpour
arXiv Machine Learning
Jun 11

Apertus LLM Family Expansion via Distillation and Quantization

arXiv:2605. 29128v2 Announce Type: replace Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating the need for the models to satisfy certain budget and hardware constraints.

By Andrei Panferov, Davit Melikidze, Martin Jaggi, Dan Alistarh
arXiv AI
Jun 12

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

arXiv:2505. 04021v3 Announce Type: replace-cross Abstract: Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increasingly important as token prices fall.

By Shan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li, Shuo Yang, Xinyuan Tong, Yang Wang, Zhiqiang Xie, Yuwei An, Shiyi Cao, Ke Bao, Deepak Vij, Xiaoning Ding, Yichen Wang, Qingda Lu, Zhong Wang, Gao Gao, Harry Xu, Junyi Shu, Jiarong Xing, Ying Sheng
arXiv AI
Aug 11

ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB

arXiv:2608. 07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge.

By Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen
arXiv Machine Learning
Sep 3

Gradient Prediction with Control Variates in the Cheap-Forward Regime

The paper investigates whether idle inference resources can help cut the high cost of scarce GPU usage during training. Using a simulated compute ledger that bills fleet work at a fraction of a GPU forward pass, the authors propose an algorithm that predicts gradients with a low‑precision, inference‑style reverse‑mode program and then refines these predictions with a few exact gradients via a control variate, turning approximation error into variance rather than bias. Experiments on a 124‑million‑parameter language model and across models ranging from 10 M to 774 M parameters show that the method can reduce simulated ledger cost when fleet work is cheap, though it also exhibits both successful transfers and failures, and does not evaluate inference‑only hardware or full optimizer‑by‑batch‑size sweeps.

By Kamil Ciosek, Nicol\`o Felicioni, Juan Elenter, Ehsan Imani
arXiv AI
Jun 10

A Theory of Training Profit-Optimal LLMs

arXiv:2605. 16430v2 Announce Type: replace-cross Abstract: Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure.

By Sophie Hao, William Merrill