arXiv:2609.24921v1 Announce Type: new
Abstract: Scientific weak signals are early, low-visibility research directions that later become central to mature scientific topics, yet existing resources suc...
By Xiao Zhou, Yilun Zhao, Owen Jiang, Tiansheng Hu, Cai Xu, Manasi Patwardhan, Arman Cohan
arXiv:2602. 20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow?
By Anirudh Ajith, Amanpreet Singh, Jay DeYoung, Nadav Kunievsky, Austin C. Kozlowski, Oyvind Tafjord, James Evans, Daniel S. Weld, Tom Hope, Doug Downey
arXiv:2608. 16645v1 Announce Type: new Abstract: Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography?
By Shaolong Chen, Yanlin Fei, Nazhou Liu, Xinmiao Yu, Lei Li, Rahul Thapa, Madalina Ciobanu, Qingqing Mao, Ritankar Das
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea.
arXiv:2606. 00644v1 Announce Type: new Abstract: AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned.
By Qiuyu Tian, Zequn Liu, Yingce Xia, Haojie Yin, Youyong Kong
arXiv:2604. 12243v2 Announce Type: replace-cross Abstract: Identifying promising research directions in fast-moving subareas is one of the most cognitively expensive tasks in modern AI research.
By Jinkai Tao, Yubo Wang, Xiaoyu Liu, Menglin Yang
arXiv:2609.22174v1 Announce Type: cross
Abstract: Anticipating emerging research directions is a critical goal of AI-assisted science. Existing methods mainly predict which concepts will co-occur in...
By Jingze Wang, Fred Sun, Shangqi Guo
arXiv:2606.22342v2 Announce Type: replace
Abstract: How does research evolve, and can we trace it at the level of individual claims? Scientific progress is not simply a uniform accumulation of facts....
By Abdul Muntakim, Md Abdullah Al Hafiz Khan, Sadid Hasan, Yong Pei
arXiv:2609.14248v1 Announce Type: cross
Abstract: Faithful citation attribution begins with identifying the intended source for a scientific claim. We study this source-identification capability thro...
By Yee Man Choi, Xuehang Guo, Songcheng Cai, Yimu Wang, Yi R. Fung, Qingyun Wang
arXiv:2609.10092v1 Announce Type: cross
Abstract: Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate be...
By Yingqian Wu, Jingcong Liang, Siyuan Wang, Zhenfei Yin, Philip Torr, Junchi Yu, Zhongyu Wei
The study investigates whether documents initially classified as noise in embedding-based topic models can be identified as precursors to emerging topics. By labeling documents based on their future trajectories and measuring confidence across multiple embedding models, the authors find that anticipatory outliers are predictable at publication time, achieving an F1 score above 0.90 on high-consensus subsets and 0.76–0.80 in chronological evaluation. The predictive power largely stems from geometric features that capture each outlier’s position in embedding space.
By Evangelia Zve, Gauvain Bourgne, Jean-Gabriel Ganascia
The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.
By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan