arXiv:2606. 08036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence.
By Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma, Hao Tian, Siqi Zhou, Wenjing Gong, Kaili Zhang, Bingqian Chen, Mitch Zhang, Yifan Yang
arXiv:2609.00747v1 Announce Type: new
Abstract: Large language models increasingly generate research ideas, yet judging their novelty or feasibility at generation time does not establish whether they...
By Fenghai Li, Zihan Tang, Haofei Yu, Yining Zhao, Jiaxuan You
arXiv:2605. 04135v2 Announce Type: replace-cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do.
By David Gringras, Misha Salahshoor
arXiv:2607. 01233v1 Announce Type: cross Abstract: LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference.
By Ziyu Chen, Yilun Zhao, Arman Cohan
arXiv:2602. 08873v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now used for academic expert recommendation.
By Lisette Esp\'in-Noboa, Gonzalo Gabriel M\'endez
arXiv:2608.31118v1 Announce Type: new
Abstract: The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled eval...
By Hamed Babaei Giglou, S\"oren Auer, Jennifer D'Souza
arXiv:2606. 24894v2 Announce Type: replace-cross Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.
By Anzhe Xie, Weihang Su, Jiaxin Mao, Yiqun Liu, Shaoping Ma, Qingyao Ai
The paper introduces EGT-KG, an evidence‑grounded typed knowledge graph retrieval framework designed to enhance scientific question answering with small language models (SLMs). It compares three QA settings—standard Retrieval‑Augmented Generation (RAG) and two EGT‑KG variants (automatically generated and expert‑defined relation schemas)—using a six‑dimensional evaluation on a biopolymer‑bound soil composite literature benchmark. Results show that both EGT‑KG variants outperform vanilla RAG, with the llama3:8b model achieving a final score of 70.37 (+14.67%) and 68.82 (+12.14%) for the AS and ES variants, respectively.
By Muran Yu, Jiechao Gao, Yuandong Pan, Barney H. Miao, Andrew C. Lesh, Kincho H. Law, Jie Wang, Michael D. Lepech
arXiv:2607. 14109v1 Announce Type: cross Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding.
By Inder Preet, Shuxin Lin, Dhaval Patel
The paper investigates whether large language models (LLMs) are narrowing the methodological diversity of archaeology. By analysing 119,000 abstracts from 2010‑2025 and running controlled experiments, the authors find only a modest shift in method use after 2023, with overall diversity actually increasing. However, LLMs tend to recommend a narrower, less diverse set of methods, especially without guidance, suggesting a potential convergence in methodological choice.
By Lorenzo Cardarelli, Roberto Ragno
arXiv:2609.01182v1 Announce Type: new
Abstract: Flagship language models appear saturated on benchmarks like MMLU (Hendrycks et al., 2021), scoring above 90% - yet benchmarks test only what the exper...
By Muhammed Saeed, Simon Razniewski
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
By Yu Fu, Yongqi Kang, Yong Zhao