arXiv:2608.30316v1 Announce Type: new
Abstract: Existing class-incremental learning methods struggle in multi-label scenarios (MLCIL) due to the inherent contradiction of learning objectives arising...
By Aoting Zhang, Dongbao Yang, Chang Liu, Xiaopeng Hong, Can Ma, Yu Zhou
The paper introduces Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients (NM-PPG), a method that relaxes the feature acquisition process to allow continuous, low‑variance policy gradients over the entire acquisition trajectory. It incorporates a straight‑through rollout that mimics discrete acquisitions during inference while enabling end‑to‑end training, and provides an average‑case upper bound on gradient variance to guide temperature sharpening. Experiments on synthetic and real datasets show that NM-PPG outperforms existing active feature acquisition baselines.
By Linus Aronsson, Morteza Haghir Chehreghani
arXiv:2507. 09471v4 Announce Type: replace Abstract: Continual Learning (CL) empowers AI models to continuously learn from sequential task streams.
By Lingfeng He, De Cheng, Zhiheng Ma, Huaijie Wang, Dingwen Zhang, Nannan Wang, Xinbo Gao
arXiv:2608. 09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization.
By Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li
arXiv:2405. 08921v2 Announce Type: replace Abstract: We focus on the online-based active learning (OAL) setting where an agent operates over a stream of observations and trades-off between the costly acquisition of information (labelled observations) and the cost of prediction errors.
By Maxime Heuillet, Ola Ahmad, Audrey Durand
arXiv:2601. 19810v2 Announce Type: replace-cross Abstract: Unsupervised pre-training can equip reinforcement learning agents with prior knowledge and accelerate learning in downstream tasks.
By Octavio Pappalardo