Bounded Precision-Geometry Scaling for Robust Multi-Task Learning under Loss Scale Mismatch
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning proposes a new scalarization method for multi-task learning that is invariant to the relative scales of task losses. By using a logarithmic transformation, SIMS converts the multi-objective problem into a single objective that preserves weak Pareto optimality and allows a smooth surrogate with controllable approximation error. Experiments on standard multi-task benchmarks show that SIMS consistently outperforms existing scalarization methods and achieves state‑of‑the‑art performance.
SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.
arXiv:2608. 10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining.
The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.
arXiv:2608.30044v1 Announce Type: new Abstract: Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an exp...
arXiv:2602. 15327v2 Announce Type: replace-cross Abstract: Machine learning model performance improvements tend to arise from competition and application.