arXiv:2606. 29182v1 Announce Type: new Abstract: Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides which hypotheses to test next.
By Dhruv Agarwal, Reece Adamson, Andrew McCallum, Peter Clark, Ashish Sabharwal, Bodhisattwa Prasad Majumder
arXiv:2602. 06448v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based scientific agents have accelerated scientific discovery, yet they often suffer from significant inefficiencies due to adherence to fixed initial priors.
By Yingming Pu, Tao Lin, Hongyu Chen
arXiv:2607. 03426v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions required in sequential decision-making settings.
By Jakob Hartmann, James Harvey, Jhonathan Navott, Erik Y. Wang, Luckeciano C. Melo, Flaviu Cipcigan, Cheng Zhang, Alessandro Abate
arXiv:2608. 15669v1 Announce Type: new Abstract: Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs.
By Zhongwei Yu, Yan Song, Xue Yan, Anjie Liu, Xingyu Lu, Yihang Chen, Huichi Zhou, Siyuan Guo, Luoyang Sun, Sihan Chen, Xiangning Yu, Jun Wang
arXiv:2608. 07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims.
By Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou
arXiv:2606. 11851v1 Announce Type: new Abstract: Open-ended scientific discovery asks agents to move beyond executing analyses for predefined questions.
By Jiayao Chen, Shi Liu, Linyi Yang
The paper introduces Agentic Reasoning for Tree Search (ARTS), a method that uses a reasoning language model to navigate the hypothesis‑experiment space in scientific discovery. Unlike traditional approaches that conflate hypothesis quality with execution quality and prune search logs, ARTS evaluates prior execution logs to distinguish implementation failures from poor hypotheses and selects the next hypothesis to pursue. By employing test‑time training to embed search‑tree knowledge into model weights, ARTS achieves a 15.3% relative improvement over leading algorithms on 22 benchmark tasks and enables smaller models like Qwen3‑4B to match or exceed the performance of larger closed‑source models at lower inference cost.
By Gurusha Juneja, Arnav Kumar Jain, Deepak Nathani, William Yang Wang, Xin Eric Wang
arXiv:2605.30219v2 Announce Type: replace
Abstract: Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and...
By Haoming Xu, Weihong Xu, Zongrui Li, Mengru Wang, Yunzhi Yao, Chiyu Wu, Jin Shang, Yu Gong, Shumin Deng
arXiv:2606. 04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks.
By Leonardo Bertolazzi, Katya Tentori, Raffaella Bernardi
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
By Hongqiang Lin, Pengfei Wang, Nenggan Zheng
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation.
arXiv:2609.01526v1 Announce Type: new
Abstract: Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM agents typically express scientific hypotheses in...
By Qing Zhao, Haowei Li, Weijian Deng, Pengxu Wei, Liang Lin