arXiv Machine Learning By Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan

Data-Native Global Optimization for Big Data K-means Clustering

Read the original on arXiv Machine Learning →

arXiv:2607. 15835v1 Announce Type: new Abstract: Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 7

Efficient Clustering with Quality Guardrails for LLM-based Recommender Systems at Industry Scale

The paper presents a scalable two‑stage clustering algorithm that guarantees per‑sample quality guardrails for large‑scale LLM‑based recommender systems. By first forming Mini‑batch K‑Means clusters and then greedily selecting representatives that meet user‑specified similarity and attribute constraints, the method ensures each sample inherits only relevant and safe outputs. Benchmarks show the approach runs faster and scales to millions of inputs, achieving a 50‑fold reduction in downstream LLM cost and runtime while maintaining personalization.

By Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer