arXiv AI

Learning-To-Measure: In-Context Active Feature Acquisition

arXiv:2510. 12624v2 Announce Type: replace-cross Abstract: Active feature acquisition (AFA) is a sequential decision-making problem where the goal is to improve model performance for test instances by adaptively selecting which features to acquire.

arXiv Machine Learning
4d ago

Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients

The paper introduces Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients (NM-PPG), a method that relaxes the feature acquisition process to allow continuous, low‑variance policy gradients over the entire acquisition trajectory. It incorporates a straight‑through rollout that mimics discrete acquisitions during inference while enabling end‑to‑end training, and provides an average‑case upper bound on gradient variance to guide temperature sharpening. Experiments on synthetic and real datasets show that NM-PPG outperforms existing active feature acquisition baselines.

By Linus Aronsson, Morteza Haghir Chehreghani
arXiv AI
Aug 11

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

arXiv:2608. 09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization.

By Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li
Hugging Face Trending Papers
Aug 10

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization. Existing task-valuation methods mostly rely on snapshot-based signals such as current pass rate or reward, which estimate how solvable a task is under the current policy.

arXiv AI
Jun 10

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

arXiv:2606. 10616v1 Announce Type: new Abstract: Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts that exceed their finite context windows, making memory retention a fundamental resource-allocation problem.

By Qingcan Kang, Liu Mingyang, Shixiong Kai, Kaichao Liang, Tao Zhong, Mingxuan Yuan