arXiv:2606. 16045v1 Announce Type: new Abstract: In the data selection problem, the objective is to choose a small, representative subset of data that can be used to efficiently train a machine learning model.
By Vincent Cohen-Addad, Sasidhar Kunapuli, Vahab Mirrokni, Mahdi Nikdan, David P. Woodruff, Samson Zhou
arXiv:2505. 03509v3 Announce Type: replace Abstract: Anomaly detection in large datasets is essential in astronomy and computer vision.
By Pablo G\'omez, Laslo E. Ruhberg, Maria Teresa Nardone, David O'Ryan
arXiv:2606. 10125v1 Announce Type: cross Abstract: Few-shot example retrieval is the dominant paradigm for grounding large language models (LLMs) in domain-specific text-to-SQL systems.
By Arash Pourhabib
arXiv:2608. 09687v1 Announce Type: new Abstract: Federated learning (FL) in Low Earth Orbit (LEO) satellite constellations is affected by non-IID data and irregular ground-station visibility, both driven by orbital geometry.
By Satwat Bashir, Tasos Dagiuklas, Muddesar Iqbal
arXiv:2608. 03249v1 Announce Type: new Abstract: Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance.
By Ning Zhu, Xiaochuan Ma, Juntao Xu, Jingze Liang, Mengfei Zhao, An Chen, Liang-Jian Deng
arXiv:2509. 11218v2 Announce Type: replace-cross Abstract: Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification.
By Johann Schmidt, Sebastian Stober