ALICE is a foundation model that estimates mutual information (MI) without per‑distribution training. Trained only on synthetic distributions, it acts as an in‑context estimator of rectified‑flow velocity fields, producing MI via a fixed identity that integrates squared differences between joint and conditional fields. The authors validate ALICE on a challenging benchmark and demonstrate its applicability to unseen data in biology, genetics, and neuroscience, achieving performance comparable to neural estimators trained separately for each distribution.
By Giulio Franzese, Simone Rossi, Pietro Michiardi
arXiv:2606. 25188v1 Announce Type: new Abstract: Efficient uncertainty quantification (UQ) is essential for trustworthy large-scale learning.
By Kun Jin, James Harrison, Jiawei Li, Sihan Liu, Jiayi Liu, Randolph Linderman, Yuening Li, Arnab Bhadury, Sourabh Prakash Bansod, Liang Liu, Jasper Snoek
arXiv:2603.05575v2 Announce Type: replace-cross
Abstract: We study prediction-powered conditional inference in the setting where labeled data are scarce, unlabeled covariates are abundant, and a blac...
By Yang Sui, Jin Zhou, Hua Zhou, Xiaowu Dai
arXiv:2606. 00241v1 Announce Type: cross Abstract: Measuring statistical dependency between high-dimensional random variables is a fundamental task in data science and machine learning.
By Zhengyang Hu, Yanzhi Chen, Hanxiang Ren, Qunsong Zeng, Youyi Zheng, Adrian Weller, Kaibin Huang, Yanchao Yang
arXiv:2605. 04847v2 Announce Type: replace-cross Abstract: Uncertainty quantification (UQ) in graph neural networks (GNNs) is crucial in high-stakes domains but remains a significant challenge.
By Soyoung park, Hwanjun Song, Sungsu Lim
The paper introduces DUO, a framework for Deep Imbalanced Regression that models each prediction as a conditional Gaussian to capture instance‑level uncertainty. By decoupling mean and variance optimization, DUO enhances learning signals for tail samples and mitigates gradient coupling that hampers hard examples. A distribution‑guided contrastive learning component further refines feature representations, leading to state‑of‑the‑art performance on several visual and biological regression benchmarks.
By Juncheng Zhou, Jiaxi Lu, Weijing Zeng, Zhong Li, Hao Qi, Jingsong Cui
arXiv:2607. 12730v1 Announce Type: cross Abstract: Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hourly, feature-limited inputs.
By Sarah Al-Shareeda, Gulcihan Ozdemir, Heung Seok Jeon
arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv:2607. 27710v1 Announce Type: new Abstract: Mutual information is a general measure of statistical dependence that captures both linear and nonlinear relationships between random variables.
By Petra Eerikinharju, Marko Tuononen, Ville Hautam\"aki
arXiv:2510.17072v2 Announce Type: replace
Abstract: Regression with non-Euclidean responses---e.g., probability distributions, networks, symmetric positive-definite matrices, and compositions---has b...
By Kyum Kim, Yaqing Chen, Paromita Dubey
arXiv:2602. 13362v2 Announce Type: replace-cross Abstract: A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty.
By \'Ad\'am Jung, Domokos M. Kelen, Andr\'as A. Bencz\'ur
arXiv:2411.02771v3 Announce Type: replace-cross
Abstract: Doubly robust estimators are widely used for estimating average treatment effects and other linear summaries of regression functions. While c...
By Lars van der Laan, Alex Luedtke, Marco Carone