Same Loss, Different Gradients
arXiv:2609.38786v1 Announce Type: new Abstract: Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the bac...
arXiv:2609.38786v1 Announce Type: new Abstract: Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the bac...
arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.
The paper investigates the difference between cost‑agnostic and cost‑sensitive loss functions when model capacity is limited. It shows that, unlike in ideal infinite‑capacity settings, optimizing a cost‑sensitive objective can yield a strictly better downstream decision than post‑processing a cost‑agnostic model. The authors prove this gap under a hypothesis class that can recover the optimal decision boundary but not the optimal cost‑agnostic hypothesis, and provide a simple example and empirical evidence on UCI datasets with simple models.
arXiv:2608. 10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining.
arXiv:2609. 07997v1 Announce Type: new Abstract: We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025).
arXiv:2609.38784v1 Announce Type: new Abstract: In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold...
arXiv:2607.22258v2 Announce Type: replace Abstract: Models trained on long-tailed data using standard softmax tend to exhibit higher training error and a larger generalisation gap for classes with fe...
arXiv:2511. 22823v2 Announce Type: replace-cross Abstract: Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly or infeasible to acquire.
arXiv:2410.18321v3 Announce Type: replace Abstract: Confidence calibration matters wherever a classifier's probabilities, not just its labels, are consumed downstream. We study Focal Calibration Loss...
arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.
arXiv:2607. 15702v2 Announce Type: replace-cross Abstract: We develop a non-asymptotic approximation, sampling, and finite-iteration optimization theory for variational physics-informed approximation of uniformly monotone nonlinear multiscale elliptic equations.
arXiv:2606. 28835v1 Announce Type: cross Abstract: Federated Learning (FL) emerged as a promising distributed machine learning paradigm.