arXiv Machine Learning By Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer

Efficient Clustering with Provable Guardrails for LLM Inference at Scale

Read the original on arXiv Machine Learning →

arXiv:2607. 19704v1 Announce Type: new Abstract: Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 7

Efficient Clustering with Quality Guardrails for LLM-based Recommender Systems at Industry Scale

The paper presents a scalable two‑stage clustering algorithm that guarantees per‑sample quality guardrails for large‑scale LLM‑based recommender systems. By first forming Mini‑batch K‑Means clusters and then greedily selecting representatives that meet user‑specified similarity and attribute constraints, the method ensures each sample inherits only relevant and safe outputs. Benchmarks show the approach runs faster and scales to millions of inputs, achieving a 50‑fold reduction in downstream LLM cost and runtime while maintaining personalization.

By Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer