arXiv Machine Learning

How Many Samples Are Enough for Learning Across Domains?

The paper investigates how many data samples per domain are needed for effective learning across multiple domains. It derives criteria from learning bounds that reveal an inverse linear relationship between the number of training domains and the required samples per domain, offering theoretical guidance for dataset adequacy and construction. The study also establishes a close link between in-domain learning and out-of-domain generalization through new generalization bounds.

arXiv Machine Learning
Jul 21

Hierarchical Domain Generalization

arXiv:2607. 16528v1 Announce Type: new Abstract: We study hierarchical domain generalization as a problem of extrapolation from finite observed regions to an entire instance space, replacing i.

By Chenxiao Yang, Zhiyuan Li, Shai Ben-David, Nathan Srebro
arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto
arXiv Machine Learning
Sep 21

Sparse Priors for Efficient Distribution Learning

arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.

By Saumya Goyal, Barnab\'as P\'oczos
arXiv Machine Learning
Sep 24

A Discrepancy-Based Perspective on Dataset Condensation

The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.

By Tong Chen, Raghavendra Selvan
arXiv Machine Learning
Jul 9

Any-Dimensional Learning by Sampling

arXiv:2607. 07680v1 Announce Type: cross Abstract: Many machine learning models are defined for inputs of different sizes, such as point clouds containing different numbers of points, sequences of tokens of different lengths, and graphs on different numbers of nodes.

By Eitan Levin, Venkat Chandrasekaran
arXiv Machine Learning
Aug 26

Joint Distribution Alignment for Universal Domain Adaptation

The paper introduces Joint Distribution Alignment for Universal Domain Adaptation (JAUA), a new algorithm designed for scenarios where source and target domains have differing label spaces. It provides a theoretical upper bound on generalization error for Universal Domain Adaptation and proposes aligning joint distributions using Chi‑Square divergence, complemented by a progressive pseudo‑labeling strategy. Experiments on six public image datasets show JAUA outperforms existing methods in handling Universal Domain Adaptation challenges.

By Shizhe Li, Hongshan Pu, Mengying Xie, Yi Xiang, Xiaowei Yang
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu