The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.
By Valentin Romanov, Monique Bax, Steven Niederer
The paper evaluates browser-based large language models (LLMs) for extracting detailed, contextualized data from scientific papers. It presents four workflows: (1) expert-curated prompts yield good extraction but struggle with nuance; (2) LLMs can generate effective prompts from simple instructions; (3) autonomous literature discovery is challenging, with missing or hallucinated references; (4) LLMs can build new datasets from guidelines that align closely with human experts, yet still need human oversight. The study outlines a practical, auditable workflow where experts set standards, models cross-check extractions, and researchers resolve disputes, enabling scalable scientific data curation.
arXiv:2604.01452v2 Announce Type: replace
Abstract: Scientific discovery is slowed by fragmented literature that requires excessive human effort to gather, analyze, and understand. AI tools, includin...
By Maxwell J. Jacobson, Daniel Xie, Jackson Shen, Adil Wazeer, Guang Lin, Xiao-Ying Yu, Haiyan Wang, Xinghang Zhang, Yexiang Xue
arXiv:2609.17291v1 Announce Type: new
Abstract: The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels....
By Marco Luca Sbodio, Marcos Mart\'inez Galindo, Vanessa Lopez, Blanca Biel, Pablo Canca, Pedro Delgado, Jes\'us I. Mendieta-Moreno, Raphael Tack, Maria J. Caturla
arXiv:2608.29612v1 Announce Type: new
Abstract: Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this...
By Shi-Ju Ran, Kun Zhang, Xi Wu, Liu-Si Yang, Wen-Jun Li
ConvergeWriter introduces a bottom‑up, data‑driven framework for long‑form document generation that first retrieves exhaustive knowledge from a source corpus and clusters it into distinct knowledge groups. These clusters then guide the creation of a hierarchical outline and the final text, ensuring the output is strictly grounded in the retrieved material and traceable to its sources. Experiments on 14B and 32B LLMs show that this approach matches or surpasses state‑of‑the‑art baselines, especially in scenarios requiring high factual fidelity and structural coherence.
By Binquan Ji, Jiaqi Wang, Ruiting Li, Xingchen Han, Yiyang Qi, Shichao Wang, Yifei Lu, Yuantao Han, Feiliang Ren