arXiv:2506. 16704v3 Announce Type: replace Abstract: We study a fundamental question of domain generalization: given a family of domains (i.
By Cynthia Dwork, Lunjia Hu, Han Shao
arXiv:2609.39512v1 Announce Type: new
Abstract: The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation a...
By Hong Zheng
arXiv:2607. 16528v1 Announce Type: new Abstract: We study hierarchical domain generalization as a problem of extrapolation from finite observed regions to an entire instance space, replacing i.
By Chenxiao Yang, Zhiyuan Li, Shai Ben-David, Nathan Srebro
arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.
By Diego Marcondes, Cl\'audia Peixoto
arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.
By Saumya Goyal, Barnab\'as P\'oczos
arXiv:2407. 12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable.
By Hong Jun Jeon, Benjamin Van Roy
The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.
By Tong Chen, Raghavendra Selvan
arXiv:2303. 18031v2 Announce Type: replace-cross Abstract: In real-world applications, a machine learning model is required to handle an open-set recognition (OSR), where unknown classes appear during the inference, in addition to a domain shift, where the data distribution differs between the training and inference phases.
By Masashi Noguchi, Shinichi Shirakawa
arXiv:2607. 07680v1 Announce Type: cross Abstract: Many machine learning models are defined for inputs of different sizes, such as point clouds containing different numbers of points, sequences of tokens of different lengths, and graphs on different numbers of nodes.
By Eitan Levin, Venkat Chandrasekaran
arXiv:2606. 23758v1 Announce Type: cross Abstract: Domain generalization learns from multiple source domains to generalize to unseen target domains.
By Xiran Wang, Jian Zhang, Lei Qi, Yang Gao, Yinghuan Shi
The paper introduces Joint Distribution Alignment for Universal Domain Adaptation (JAUA), a new algorithm designed for scenarios where source and target domains have differing label spaces. It provides a theoretical upper bound on generalization error for Universal Domain Adaptation and proposes aligning joint distributions using Chi‑Square divergence, complemented by a progressive pseudo‑labeling strategy. Experiments on six public image datasets show JAUA outperforms existing methods in handling Universal Domain Adaptation challenges.
By Shizhe Li, Hongshan Pu, Mengying Xie, Yi Xiang, Xiaowei Yang
The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.
By Shaojie Li, Yunbei Xu