arXiv:2608. 10441v1 Announce Type: new Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using.
By Ying Yuan
arXiv:2407. 12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable.
By Hong Jun Jeon, Benjamin Van Roy
arXiv:2606. 19476v1 Announce Type: cross Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect.
By Eric Elmoznino, Sangnie Bhardwaj, Johannes von Oswald, Rajai Nasser, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento, Rif A. Saurous, Guillaume Lajoie
The paper introduces TRACE, a digital‑advertising diagnostic environment that uses simulated interventions to generate verifiable rewards for training reasoning agents. By injecting controlled interventions into a simulator, the hidden cause of anomalies becomes an oracle label, enabling agents to learn to identify root causes and affected segments through noisy, confounded evidence. Experiments show that reinforcement learning with these synthesized rewards outperforms large prompted baselines, achieving higher accuracy while using fewer tool calls.
By Rui Sun, Zhan Shi, Bing He
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge.
The paper proposes treating the ‘unit’—a persistent referent that multiple events may refer to—as an explicit primitive in machine learning tasks. It formalizes supervised learning as learning a pair of a tokenizer that generates a contextual unit token and a shared response law that uses this token, thereby distinguishing homogeneous from heterogeneous worlds. The work also introduces concepts such as unit abduction and trusted resolvers to handle cases where unit identity is unresolved.
By Heyang Gong
arXiv:2603. 12109v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a de facto paradigm for building LLM-based agents that act, interact, and reason over extended task horizons.
By Deyu Zou, Yongqiang Chen, Fan Feng, Mufei Li, Pan Li, Yu Gong, James Cheng
The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.
By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
arXiv:2605.14040v2 Announce Type: replace
Abstract: Trackable improvement in multimodal physics reasoning rests on a training-and-evaluation system that is itself rarely verified: the corpora a model...
By Shan Yang
arXiv:2607. 07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn?
By Anne Harrington, Nayan Saxena, Michael Murphy, Anastasia Borovykh, Zeyu Yun, Sridhar Kamath, Ara Eindra Kyi, Trevor Darrell, Jitendra Malik, Yutong Bai
The paper "Demystifying Reinforcement Learning Post-Training of Language Models" investigates how reinforcement learning (RL) post‑training enhances large language models (LLMs) for tasks such as reasoning, math, and coding. By isolating RL components in a controlled setting, the authors analyze how the base model’s prior distribution, reward granularity, prompt diversity, and model scale influence outcomes, using policy entropy to compare pre‑training, supervised fine‑tuning (SFT), and RL stages. The study clarifies the role of spurious rewards, the importance of the base model’s probability mass on desired behaviors, and how these factors interact to determine post‑training success, offering a practical primer for NLP researchers.
"whyItMatters":"The work provides a clearer understanding of RL post‑training mechanics, helping researchers and practitioners effectively apply RL to improve LLM capabilities."
By Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh, Natasha Jaques
The paper introduces Transfer Learning for Evolving Domains (TrED), a framework that models how data availability changes over time in real-world applications. TrED treats the entire trajectory of model updates as a single learning problem, rather than isolated snapshots, and defines a data availability process, a flexible learning protocol, and an evaluation criterion that scores the whole trajectory. The authors review existing transfer learning methods, noting that most are tailored to specific regimes and do not optimize the full trajectory, and argue that TrED is a well‑posed, unsolved research direction.
By Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Pedro Ribeiro, Pedro Saleiro, Pedro Bizarro, Carlos Soares