SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning proposes a new scalarization method for multi-task learning that is invariant to the relative scales of task losses. By using a logarithmic transformation, SIMS converts the multi-objective problem into a single objective that preserves weak Pareto optimality and allows a smooth surrogate with controllable approximation error. Experiments on standard multi-task benchmarks show that SIMS consistently outperforms existing scalarization methods and achieves state‑of‑the‑art performance.
By Zebin Chen, Fei Xing, Yang Chen, Hua Liu, Andy HF Chow, Yuhua Qian, Yu Zhang
SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.
By Liang You, Hengyu Shi, Dongwen Ou
arXiv:2608. 10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining.
By Lening Zhao, Qipeng Zhan, Li Shen
The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.
By Jhen-Ke Lin, Hong-Yun Lin
arXiv:2608.30044v1 Announce Type: new
Abstract: Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an exp...
By Jhen-Ke Lin
arXiv:2602. 15327v2 Announce Type: replace-cross Abstract: Machine learning model performance improvements tend to arise from competition and application.
By Hanlin Zhang, Jikai Jin, Vasilis Syrgkanis, Sham Kakade
arXiv:2608. 20061v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost.
By Nayeon Kim, Hojin Lee, Yunju Bak, Jaesun Park, Boseop Kim
arXiv:2606. 10632v1 Announce Type: cross Abstract: Lipschitz-style individual fairness formalizes the idea that semantically similar examples should receive similar predictions, but its evaluation in multi-task learning (MTL) can be confounded by method-induced representation scales.
By Junbo Ding, Xin Zang, Chenchen Pan, Donghao Song, Jiaxin Zhu, Danhuai Guo
arXiv:2602. 03001v2 Announce Type: replace-cross Abstract: To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and costly to tune.
By Hiroki Naganuma, Shagun Gupta, Youssef Briki, Ioannis Mitliagkas, Irina Rish, Parameswaran Raman, Hao-Jun Michael Shi
arXiv:2608. 09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predictable changes in model function.
By Chencheng Zhu, Xiaoyang Li, Taotao Cai
arXiv:2606. 31591v1 Announce Type: cross Abstract: Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned behaviour on unrelated prompts.
By Jason R. Brown, Patrick Leask, Lev McKinney
arXiv:2607. 23860v1 Announce Type: new Abstract: Deep ensembles provide the most reliable uncertainty estimates in deep learning, but their cost grows linearly with the number of members.
By Mihai Suteu, Ovidiu Serban