arXiv AI By Prayas Agrawal, Prateek Chanda, Ishita Khatri, Ganesh Ramakrishnan, Bamdev Mishra, Pratik Jawanpuria

Minibatch Selection via Partition Matroid Constrained Gradient Matching

Read the original on arXiv AI →

arXiv:2606. 07954v1 Announce Type: cross Abstract: Training large language models (LLMs) on heterogeneous data requires selecting minibatches that balance convergence speed with coverage across domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

The paper introduces TESS, a scalable data‑selection framework that replaces per‑sample weights with a selection network to improve transferability across datasets and model sizes. It identifies instability in existing meta‑learning for training‑data selection (MTS) due to weight suppression and overreliance on easy features, and proposes a Pointwise Value Matching objective to address these issues. Experiments on large language model safety and instruction tuning show strong transfer from subsets to full corpora and from smaller to larger models.

By Zilin Du, Bowen Yang, Boyang Albert Li