arXiv AI By Mansi Sakarvadia, Marco Ciccone, Colin Raffel

Initialization Improves LLM-Driven Discovery

Read the original on arXiv AI →

The paper investigates how the set of prior iterates influences success in large language model (LLM)-driven discovery tasks. It introduces 12 new harnesses called 'Modular' and evaluates them on five diverse discovery problems, revealing that success is fragile and highly dependent on harness design. The study identifies mode collapse—a sharp loss of iterate diversity—as a common failure, and shows that early discoveries predict final outcomes, leading to a new initialization strategy that consistently improves performance across harnesses and applications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 10

Towards Diverse Scientific Hypothesis Search with Large Language Models

arXiv:2606. 10587v1 Announce Type: cross Abstract: Large language models (LLMs) are on the rise for accelerating scientific discovery, most recently in advanced tasks such as generating valid scientific hypotheses.

By Haorui Wang, Parshin Shojaee, Kazem Meidani, Kunyang Sun, Jos\'e Miguel Hern\'andez-Lobato, Teresa Head-Gordon, Jiajun He, Chandan K. Reddy, Chao Zhang, Yuanqi Du
arXiv AI
Jul 29

Structured Scaling of AI Discovery Across Diverse Scientific Domains

arXiv:2604. 19341v2 Announce Type: replace-cross Abstract: Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions.

By Haotian Ye, Haowei Lin, Jingyi Tang, Yizhen Luo, Rahul Thapa, Caiyin Yang, Chang Su, Rui Yang, Ruihua Liu, Rundao Li, Zeyu Li, Pengwei Sun, Chong Gao, Dachao Ding, Guangrong He, Miaolei Zhang, Lina Sun, Wenyang Wang, Yuchen Zhong, Zhuohao Shen, Puheng Li, Pan Lu, Bianxiao Cui, Di He, Jianzhu Ma, Junfeng Li, Hexi Baoyin, Yejin Choi, Stefano Ermon, Xiaowen Chu, Tongyang Li, Yuzhi Xu, James Zou