The paper introduces an agentic forecasting environment built on 2,100+ resolved Polymarket questions, where a language model (Qwen3.5-35B-A3B) learns to gather evidence during rollout via web search, page reading, and financial time series, all filtered to avoid post‑cutoff leaks. Training with single‑epoch GRPO and a Brier‑score reward improves calibration by 30‑40% and reduces search attempts, while the trained policy outperforms four frontier models in evidence‑based forecasting, achieving lower soft‑Brier scores at roughly 5% of the inference cost. The authors release the environment, dataset, and per‑rollout records as a reusable harness for temporal forecasting agents.
By Yusuf Afifi, Artur Kiulian, Anton Polishko, Mykola Khandoga, Hamudi Naanaa, Alina Krasnobrizha
EvolveTrade is a self‑evolving framework that treats the system prompt of a tool‑using LLM trading agent as a text‑parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and portfolio feedback while keeping the backbone LLM fixed, allowing the agent to refine its information‑acquisition and portfolio‑construction procedures over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed‑policy baselines, with behavioral analyses indicating increased code‑mediated analysis and regime‑relevant computations.
By Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang
CAST is a critique‑aware training framework that transforms sparse task outcomes into action‑level supervision for both critique learning and policy optimization. By analyzing agent trajectories, CAST synthesizes structured rationales that explain action validity under partial observability, enabling the creation of richer training data. Fine‑tuned Qwen3‑family models trained with CAST show significant reliability gains, outperforming GPT‑OSS‑120B by over 10% on Retail tasks and improving Telehealth performance by 9% in an out‑of‑domain setting.
By Amir Saeidi, Zehua Zhang, Rishitosh Singh, Naman Ahuja, Vivek Gupta, Ali Payani, Gaowen Liu, Jayanth Srinivasa, Chitta Baral
arXiv:2609.05435v2 Announce Type: replace
Abstract: Can language agents continually learn from experience, turning earlier interactions into reusable capabilities? AhaBench evaluates this ability thr...
By Zerui Cheng, Jiawei Xu, Huacan Chai, Jiayang Sun, Pramod Viswanath, Maxm Pan
arXiv:2606. 02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
By Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su
arXiv:2606. 02461v2 Announce Type: replace Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
By Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su
arXiv:2607. 21419v1 Announce Type: new Abstract: In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization.
By Yipeng Shi, Zhipeng Ma, Yue Wang, Qitai Tan, Yang Li, Peng Chen, Zhengzhou Zhu
arXiv:2607. 07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn?
By Anne Harrington, Nayan Saxena, Michael Murphy, Anastasia Borovykh, Zeyu Yun, Sridhar Kamath, Ara Eindra Kyi, Trevor Darrell, Jitendra Malik, Yutong Bai
FinSkillBench is an evaluation suite that tests whether language model agents can use financial domain skills to solve investment management tasks across portfolio construction, risk management, and fundamental analysis. The benchmark contains 12 subtasks with 2,603 episodes, each providing point‑in‑time inputs, hidden ground truth, and a verifier. Experiments show that curated skill packages improve performance significantly, while self‑generated skills offer little benefit, indicating that reliable procedural skills are crucial for effective AI agents in this domain.
By Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun
arXiv:2609.24862v1 Announce Type: new
Abstract: Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning s...
By Yifan Hu, Xilin Dai, Zhiyuan Qu, Yiding Liu, Zewei Dong, Jiang-ming Yang, Qiang Xu
arXiv:2606. 00143v1 Announce Type: cross Abstract: Financial markets are inherently non-stationary, exhibiting frequent regime shifts and structural changes that render traditional Portfolio Management (PM) approaches ineffective.
By Chaofan Pan, Lingfei Ren, Linbo Xiong, Yonghao Li, Wei Wei, Xin Yang
arXiv:2605.08693v3 Announce Type: replace
Abstract: Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and...
By Min Yang, Jinghua Piao, Xu Xia, Xiaochong Lan, Jiaju Chen, Yongshun Gong, Yong Li