arXiv:2607. 05456v1 Announce Type: new Abstract: While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication.
By Ramsha Kamran, Maheera Amjad, Zartasha Mustansar, Arsalan Shaukat, Salma Sherbaz, Muhammad U. S. Khan
arXiv:2609.22235v1 Announce Type: new
Abstract: While multi-agent systems based on large language models (LLMs) have shown promise in automating the progressive workflow of academic research, extendi...
By Yuhe Wu, Guangyu Wang, Jiaxin Liu, Guang Zhang
arXiv:2605.30947v4 Announce Type: replace
Abstract: LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and qu...
By Yating Pan, Jiajun Zhang, Jun Wang, Qi Su
DataSTORM is an LLM‑based agentic system designed to conduct deep research over large‑scale structured databases and internet sources. It applies principles of Exploratory Data Analysis and Data Storytelling to frame research as a thesis‑driven analytical process, iteratively generating hypotheses, performing quantitative reasoning, and crafting coherent narratives. Evaluations on InsightBench and a new ACLED‑based dataset show that DataSTORM surpasses existing systems, achieving significant improvements in insight‑level recall and summary‑level scores.
By Shicheng Liu, Yucheng Jiang, Sajid Farook, Camila Nicollier Sanchez, David Fernando Castro Pena, Monica S. Lam
arXiv:2609.22104v1 Announce Type: new
Abstract: As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottle...
By Rongcan Pei, Fang Guo, Qinglin Qi, Qi Zhu, Yun Luo, Jianhao Yan, Minjun Zhu, Qiujie Xie, Dehong Zheng, Yue Zhang
DeepWeaver is a framework designed to improve open‑ended question answering by weaving noisy retrieved evidence into comprehensive, well‑cited answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and supporting evidence, and uses subordinate TBCs to refine and expand the evidence before final generation. Evaluations on LoQA and DeepResearch Bench show that DeepWeaver enhances content sufficiency, citation grounding, and detail preservation across multiple LLMs.
By Xujia Wang, Yizhe Zhang, Bin Xu, Lei Hou, Juanzi Li
DeepWeaver addresses the evidence synthesis gap in open‑ended question answering by weaving noisy retrieved evidence into comprehensive answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and citations, allowing the system to revise and expand evidence before final generation. Evaluations on LoQA and DeepResearch Bench show improved content sufficiency, citation grounding, and detail preservation across multiple LLMs.
Deep Research Bench II is a new benchmark designed to evaluate Deep Research Agents (DRAs) by requiring them to produce research reports for 132 grounded tasks across 22 domains. Each report is assessed using 9,430 fine‑grained binary rubrics that cover information recall, analysis, and presentation, all derived from expert‑written investigative articles through a rigorous LLM‑plus‑human pipeline. Evaluation of current state‑of‑the‑art DRAs shows that even the best models satisfy fewer than 50% of these rubrics, highlighting a significant gap between automated agents and human experts.
By Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao
arXiv:2606. 01613v1 Announce Type: cross Abstract: This paper presents an agentic retrieval-augmented generation (RAG) framework for domain-specific technical reasoning support, instantiated over a curated corpus of approximately 2,100 academic papers in intelligent tires, vehicle dynamics, and vehicle control.
By Kanwar Bharat Singh
arXiv:2508.14273v3 Announce Type: replace
Abstract: As researchers increasingly adopt LLMs as writing assistants, generating high-quality research paper introductions remains both challenging and ess...
By Krishna Garg, Firoz Shaik, Sambaran Bandyopadhyay, Cornelia Caragea
arXiv:2607. 20498v1 Announce Type: new Abstract: Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks.
By Fanjin Zhang, Zhengyang Wang, Ruixuan Huang, Kefan Zhang, Amy Xin, Yuanchun Wang, Shu Zhao, Evgeny Kharlamov, Jie Tang, Juanzi Li
arXiv:2602. 20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow?
By Anirudh Ajith, Amanpreet Singh, Jay DeYoung, Nadav Kunievsky, Austin C. Kozlowski, Oyvind Tafjord, James Evans, Daniel S. Weld, Tom Hope, Doug Downey