arXiv AI By Prayas Agrawal, Prateek Chanda, Ishita Khatri, Ganesh Ramakrishnan, Bamdev Mishra, Pratik Jawanpuria

Minibatch Selection via Partition Matroid Constrained Gradient Matching

Read the original on arXiv AI →

arXiv:2606. 07954v1 Announce Type: cross Abstract: Training large language models (LLMs) on heterogeneous data requires selecting minibatches that balance convergence speed with coverage across domains.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.