arXiv:2506. 16704v3 Announce Type: replace Abstract: We study a fundamental question of domain generalization: given a family of domains (i.
By Cynthia Dwork, Lunjia Hu, Han Shao
The paper investigates how many data samples per domain are needed for effective learning across multiple domains. It derives criteria from learning bounds that reveal an inverse linear relationship between the number of training domains and the required samples per domain, offering theoretical guidance for dataset adequacy and construction. The study also establishes a close link between in-domain learning and out-of-domain generalization through new generalization bounds.
By Hong Zheng
arXiv:2609.39512v1 Announce Type: new
Abstract: The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation a...
By Hong Zheng
arXiv:2303. 18031v2 Announce Type: replace-cross Abstract: In real-world applications, a machine learning model is required to handle an open-set recognition (OSR), where unknown classes appear during the inference, in addition to a domain shift, where the data distribution differs between the training and inference phases.
By Masashi Noguchi, Shinichi Shirakawa
arXiv:2606. 23758v1 Announce Type: cross Abstract: Domain generalization learns from multiple source domains to generalize to unseen target domains.
By Xiran Wang, Jian Zhang, Lei Qi, Yang Gao, Yinghuan Shi
The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.
By Tong Chen, Raghavendra Selvan