The paper investigates how many data samples per domain are needed for effective learning across multiple domains. It derives criteria from learning bounds that reveal an inverse linear relationship between the number of training domains and the required samples per domain, offering theoretical guidance for dataset adequacy and construction. The study also establishes a close link between in-domain learning and out-of-domain generalization through new generalization bounds.
By Hong Zheng
arXiv:2609.36030v1 Announce Type: new
Abstract: The lack of rigorous safety and performance certificates remains a key bottleneck to the deployment of modern learning-based methods. Sample compressio...
By Dario Paccagnan, Marius Tirlea
arXiv:2606. 08638v1 Announce Type: cross Abstract: Recent research has developed practical, parallelizable first-order methods for large scale linear programming, but performance is highly dependent on hyperparameter selection.
By Siddharth Prasad, Dravyansh Sharma
arXiv:2606. 17319v1 Announce Type: cross Abstract: Motivated by the optimization of bounded binary black-box functions, we study the problem of learning polynomial surrogates over the Boolean hypercube.
By Jasper van Doornmalen, Mathieu Molina, Victor Verdugo, Jos\'e Verschae
arXiv:2607. 22467v1 Announce Type: new Abstract: Data scarcity poses a fundamental challenge in training generative models to produce initial guesses for parametric optimization problems that are otherwise numerically expensive to solve.
By Anjian Li, Ryne Beeson
arXiv:2601. 20970v3 Announce Type: replace-cross Abstract: The maximum-entropy remote sampling problem (MERSP) is to select a subset of $s$ random variables from a set of $n$ random variables, so as to maximize the information concerning a set of target random variables that are not directly observable.
By Gabriel Ponte, Marcia Fampa, Jon Lee