arXiv Machine Learning

Adaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification

The paper introduces the Adaptive Margin Ordinal Loss (AMOL), a new loss function designed to reduce the tendency of neural networks to predict center classes in ordinal classification tasks—a problem called center‑class hedging. AMOL applies a multiplicative weight to per‑class loss terms that is large only when a candidate class is near the center while the true label is far from it, thereby discouraging hedging. The authors also propose the Center‑Hedging Rate (CHR) metric to quantify this failure mode and demonstrate that AMOL achieves state‑of‑the‑art Quadratic Weighted Kappa scores on four benchmarks, with an asymmetric variant eliminating hedging on the Abalone dataset.

arXiv Machine Learning
Sep 22

Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions

The paper benchmarks six long‑tail loss functions—cross‑entropy, weighted CE, class‑balanced loss, focal loss, LDAM, and logit‑adjusted softmax—across three single‑cell foundation model architectures (scGPT, scBERT, Geneformer) and three datasets (Multiple Sclerosis, Zheng68K, human Pancreas). It shows that overall accuracy masks systematic failures on rare, disease‑relevant cell types, with a consistent gap between overall accuracy, Macro‑F1, and rare‑class recall under plain cross‑entropy. The study identifies two distinct regimes of rare‑class failure, predicts reweighting efficacy by absolute training‑set size, and finds class‑balanced loss and LDAM to be the most reliable across all settings.

By Zeyu Dong, Jiahui Zhong
arXiv AI
Jun 18

Generalized Kullback-Leibler Divergence Loss

arXiv:2503. 08038v2 Announce Type: replace-cross Abstract: In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of (1) a weighted Mean Square Error (wMSE) loss and (2) a Cross-Entropy loss incorporating soft labels.

By Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong
arXiv AI
Jun 30

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.

By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
arXiv AI
4d ago

PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions

The paper introduces PACT, a tuning method for single-token typed-decision models that leverages contrastive pair data to add four training terms—difference-in-differences margin, permutation-consistency, evidence-necessity, and ordinal transport cost—without requiring new annotations. PACT achieves comparable accuracy to existing recipes while reducing position bias and ordinal error, and it improves robustness and stability across seeds. The authors provide code, data splits, and trained adapters for reproducibility.

By Yida Lin
arXiv AI
Aug 13

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

arXiv:2608. 11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.

By Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
arXiv Machine Learning
Sep 14

RiPPLE: Cross-Space Performance Prediction from Early Training for Neural Architecture Search

RiPPLE is a method for ranking neural architectures across an entire search space using only a small fraction of early training data. It treats partial training as labels for a limited set of anchor architectures, extrapolates their learning curves, and propagates these surrogate labels to other architectures without requiring per‑candidate features. The approach is evaluated on twelve benchmark cells from four search‑space families and the larger DARTS space, demonstrating its effectiveness in ranking quality, label efficiency, and architecture selection.

By Yifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li, Jiaojiao Jiang