The paper investigates how many data samples per domain are needed for effective learning across multiple domains. It derives criteria from learning bounds that reveal an inverse linear relationship between the number of training domains and the required samples per domain, offering theoretical guidance for dataset adequacy and construction. The study also establishes a close link between in-domain learning and out-of-domain generalization through new generalization bounds.
By Hong Zheng
arXiv:2607. 16528v1 Announce Type: new Abstract: We study hierarchical domain generalization as a problem of extrapolation from finite observed regions to an entire instance space, replacing i.
By Chenxiao Yang, Zhiyuan Li, Shai Ben-David, Nathan Srebro
arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.
By Saumya Goyal, Barnab\'as P\'oczos
arXiv:2606. 28309v1 Announce Type: cross Abstract: Binary classification from positive-only samples is a variant of PAC learning in which the learner receives i.
By Shai Ben-David, Farnam Mansouri, Anay Mehrotra, Manolis Zampetakis
arXiv:2608.30246v1 Announce Type: cross
Abstract: The fundamental theorem of statistical learning states that, under suitable measurability assumptions, finite Vapnik--Chervonenkis (VC) dimension gua...
By Mateus Jesus de Arruda Campos, Gabriel Fernandes, Vinicius de Oliveira Rodrigues
arXiv:2609.39512v1 Announce Type: new
Abstract: The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation a...
By Hong Zheng