Linear models with an $L_1$-norm penalty remain state-of-the-art for high-dimensional ($d > 1,000,000$) tasks, offering a straightforward method for solving real-world industry problems. Despite their...
Linear models with an $L_1$-norm penalty are still the leading approach for high‑dimensional tasks, yet many existing solvers are slow, ineffective, and hard to parallelise, making them unsuitable for large industry‑scale corpora. The paper evaluates several recent state‑of‑the‑art methods and shows that older techniques outperform them in general use. It also demonstrates that a simple baseline—LBFGS applied to a sub‑gradient with minor tweaks—yields strong performance and is easier to support and scale in production.
By Edward Raff, James Holt
Fine-tune a cost-efficient model with the outputs of a large frontier model–all on the OpenAI platform
arXiv:2603. 24963v3 Announce Type: replace Abstract: Modern computational advertising platforms typically rely on recommendation systems to predict user responses, such as click-through rates, conversion rates, and other optimization events.
By Jiang Liu, John Martabano Landy, Yao Xuan, Swamy Muddu, Nhat Le, Munaf Sahaf, Luc Kien Hang, Rupinder Khandpour, Kevin De Angeli, Chang Yang, Shouyuan Chen, Shiblee Sadik, Anirudh Agrawal, Djordje Gligorijevic, Jingzheng Qin, Peggy Yao, Alireza Vahdatpour
arXiv:2605. 29128v2 Announce Type: replace Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating the need for the models to satisfy certain budget and hardware constraints.
By Andrei Panferov, Davit Melikidze, Martin Jaggi, Dan Alistarh
arXiv:2505. 04021v3 Announce Type: replace-cross Abstract: Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increasingly important as token prices fall.
By Shan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li, Shuo Yang, Xinyuan Tong, Yang Wang, Zhiqiang Xie, Yuwei An, Shiyi Cao, Ke Bao, Deepak Vij, Xiaoning Ding, Yichen Wang, Qingda Lu, Zhong Wang, Gao Gao, Harry Xu, Junyi Shu, Jiarong Xing, Ying Sheng
arXiv:2608. 07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge.
By Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen
arXiv:2608. 13573v1 Announce Type: new Abstract: Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems.
By William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang
The paper investigates whether idle inference resources can help cut the high cost of scarce GPU usage during training. Using a simulated compute ledger that bills fleet work at a fraction of a GPU forward pass, the authors propose an algorithm that predicts gradients with a low‑precision, inference‑style reverse‑mode program and then refines these predictions with a few exact gradients via a control variate, turning approximation error into variance rather than bias. Experiments on a 124‑million‑parameter language model and across models ranging from 10 M to 774 M parameters show that the method can reduce simulated ledger cost when fleet work is cheap, though it also exhibits both successful transfers and failures, and does not evaluate inference‑only hardware or full optimizer‑by‑batch‑size sweeps.
By Kamil Ciosek, Nicol\`o Felicioni, Juan Elenter, Ehsan Imani
arXiv:2605. 16430v2 Announce Type: replace-cross Abstract: Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure.
By Sophie Hao, William Merrill