Rocket Money x Hugging Face: Scaling Volatile ML Models in Production
Related stories
Model Distillation in the API
Fine-tune a cost-efficient model with the outputs of a large frontier model–all on the OpenAI platform
GaLore: Advancing Large Model Training on Consumer-grade Hardware
Design Once, Deploy at Scale: Template-Driven ML Development for Large Model Ecosystems
arXiv:2603. 24963v3 Announce Type: replace Abstract: Modern computational advertising platforms typically rely on recommendation systems to predict user responses, such as click-through rates, conversion rates, and other optimization events.
Apertus LLM Family Expansion via Distillation and Quantization
arXiv:2605. 29128v2 Announce Type: replace Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating the need for the models to satisfy certain budget and hardware constraints.
Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
arXiv:2505. 04021v3 Announce Type: replace-cross Abstract: Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increasingly important as token prices fall.
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
arXiv:2608. 07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge.
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
arXiv:2608. 13573v1 Announce Type: new Abstract: Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems.
A Theory of Training Profit-Optimal LLMs
arXiv:2605. 16430v2 Announce Type: replace-cross Abstract: Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure.
Large Databases Need Small, Open-Weight Language Models
arXiv:2606. 31808v1 Announce Type: new Abstract: Language model systems built around proprietary APIs often operate on a token-based cost model.
Mojo: A Promising Tool for Scalable Financial AI Efficiency
arXiv:2606. 16059v1 Announce Type: cross Abstract: For thirty years, quantitative finance has paid a costly two-language tax: models researched in Python are rewritten in C++ for production, often introducing numerical discrepancies.
Towards Engineering Scaling Laws with Pretraining Data Composition
arXiv:2606. 19781v1 Announce Type: cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size.