arXiv:2606. 08574v1 Announce Type: new Abstract: Data pruning (DP), as an oft-stated strategy to alleviate heavy training burdens, reduces the volume of training samples according to a well-defined pruning method while striving for near-lossless performance.
By Chenhan Jin, Shengze Xu, Qingsong Wang, Fan Jia, Dingshuo Chen, Tieyong Zeng
arXiv:2608.30699v1 Announce Type: cross
Abstract: Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to...
By Yue Cheng, Jiajun Zhang, Xiaohui Gao, Weiwei Xing, Zhanxing Zhu
arXiv:2609.16365v1 Announce Type: cross
Abstract: Real-world datasets often exhibit long-tailed class distributions, where a few head classes contain a large number of training samples while a large...
By Siyu Yuan
arXiv:2606. 25488v1 Announce Type: new Abstract: Knowledge Distillation (KD) is widely used to obtain compact models for efficient inference in resource-constrained environments.
By Yifan Wu, Yiqi Wang, Xichen Ye, Wenjing Yan, Xiaoqiang Li, Cheng Jin, Xiangyu Yue, Weizhong Zhang
arXiv:2409. 13007v3 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge in classification tasks, often causing standard learning algorithms to become biased toward the majority class.
By Asif Newaz, Asif Ur Rahman Adib, Taskeed Jabid
arXiv:2609.10311v1 Announce Type: cross
Abstract: The lottery ticket hypothesis posits the existence of winning tickets: sparse subnetworks that, when trained in isolation from their original initial...
By Benedikt Tscheschner, Eduardo Veas, Marc Masana
arXiv:2604. 02765v2 Announce Type: replace Abstract: Class-incremental learning (CIL) is commonly evaluated under predefined schedules with fixed or nearly equal class increments, leaving irregular class-arrival scenarios underexplored.
By Zhiming Xu, Baile Xu, Jian Zhao, Furao Shen, Suorong Yang
arXiv:2506. 20893v5 Announce Type: replace-cross Abstract: In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten class.
By Ali Ebrahimpour-Boroojeny, Yian Wang, Hari Sundaram
arXiv:2603. 17205v3 Announce Type: replace-cross Abstract: Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process.
By Haoyang Fang, Shuai Zhang, Yifei Ma, Hengyi Wang, Cuixiong Hu, Katrin Kirchhoff, Bernie Wang, George Karypis
arXiv:2607. 01474v1 Announce Type: new Abstract: Class imbalance poses a critical challenge in federated learning (FL), where underrepresented classes suffer from poor predictive performance yet cannot be addressed by standard centralized techniques due to privacy and heterogeneity constraints.
By Haemin Park, Diego Klabjan, Martin W. Braun, Xiuqi Li, Balakrishnan Ananthanarayanan
CRISP (Coreset Reduction via Importance-Stratified Pruning) is a linear-time method that reduces negative-class examples in highly imbalanced tabular datasets by allocating a budget across quantile strata of a proxy-model score and using sample weights to correct for unequal inclusion probabilities. On a production fraud dataset, CRISP cuts the training set from 25 M to about 1.70 M rows (a 93.2% reduction) while preserving 99.7% of the full-data Average Precision. In public benchmarks such as CriteoPrivateAds, CRISP consistently achieves the highest mean Average Precision across a range of majority reductions, with ablation studies highlighting budget allocation and inverse-propensity weighting as key contributors to its performance.
By Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani
arXiv:2506. 01486v2 Announce Type: replace Abstract: Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability.
By Jelke Wibbeke, Sebastian Rohjans, Andreas Rauh