arXiv:2609.14065v1 Announce Type: new
Abstract: When algorithmic predictions inform people's decisions, the models we deploy are performative and actively shape the data we see. This feedback loop be...
By Gabriele Farina, Juan Carlos Perdomo
arXiv:2201. 01973v3 Announce Type: replace-cross Abstract: The problem of linear predictions has been extensively studied for the past century under pretty generalized frameworks.
By Saptarshi Chakraborty, Debolina Paul, Swagatam Das
arXiv:2606. 06855v1 Announce Type: cross Abstract: While algorithmic stability is a central tool for understanding generalization of learning algorithms, existing high-probability guarantees typically rely on uniform boundedness or sub-Gaussian/sub-Weibull tail assumptions, which can be overly restrictive for modern settings with heavy-tailed or unbounded losses.
By Qianqian Lei, Soham Bonnerjee, Yuefeng Han, Wei Biao Wu
arXiv:2609.06430v1 Announce Type: new
Abstract: We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Stati...
By Yiming Ying
arXiv:2602. 20971v3 Announce Type: replace-cross Abstract: Bubeck and Selke (2021) propose the connection between the Law of Robustness and robust generalization error as an open problem.
By Mihir More, Aritra Das, Jaee Ponde, Himadri Mandal, Vishnu Varadarajan, Debayan Gupta
arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.
By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
arXiv:2606. 22775v2 Announce Type: replace-cross Abstract: Distribution shift between training and deployment is a pervasive challenge for modern AI systems.
By Zhewen Hou, Tian Zheng
arXiv:2606. 15600v1 Announce Type: cross Abstract: Cardinality-estimation (CE) research ranks estimators by q-error, yet it is well known that q-error is an imperfect proxy for query-plan quality.
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2505. 20178v2 Announce Type: replace-cross Abstract: Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation.
By Pranav Mani, Peng Xu, Zachary C. Lipton, Michael Oberst
arXiv:2602. 17894v2 Announce Type: replace-cross Abstract: Data collection is a critical component of modern statistical and machine learning pipelines, particularly when data must be gathered from multiple heterogeneous sources to study a target population of interest.
By Michael O. Harding, Vikas Singh, Kirthevasan Kandasamy
arXiv:2606. 16883v1 Announce Type: cross Abstract: Generalization is a critical property of data-driven models, particularly deep learning models deployed in safety-critical applications.
By Abdul-Rauf Nuhu, Parham M. Kebria, Vahid Hemmati, Mahmoud N. Mahmoud, Edward Tunstel, Abdollah Homaifar
arXiv:2506. 20573v4 Announce Type: replace-cross Abstract: Public datasets, crucial for modern machine learning and statistical inference, often contain low-quality or contaminated samples that can harm model performance.
By Kristian Minchev, Dimitar I. Dimitrov, Nikola Konstantinov