REPAIR is a self‑evolving data augmentation framework designed to improve scientific dense retrievers by addressing long‑tail concept gaps and fact sensitivity. It iteratively generates training data through diagnosis of long‑tail concepts, API‑guided evidence expansion, and hard negative mining, thereby grounding retrieval in factual reality. Experiments show that REPAIR outperforms 19 strong baselines across nine materials science and biomedical benchmarks.
By Yerim Oh, Gunhee Kim
The paper introduces the task of Scientific Claim Unlearning and presents a new benchmark, SciUnlearn, to evaluate it. It highlights that language models trained on static scientific corpora risk disseminating outdated or retracted claims as scientific knowledge evolves. Current machine unlearning methods fail to effectively remove claim-level knowledge, often only suppressing it superficially, underscoring the need for specialized techniques for structured knowledge removal.
By Snigdha Paul, Manasi Patwardhan, Arman Cohan
arXiv:2607. 20926v1 Announce Type: new Abstract: Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources.
By Yinhao Tang, Youqing Fang, Yanan Sun, Wenran Liu, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen
arXiv:2606.21359v2 Announce Type: replace
Abstract: Large language models (LLMs) are increasingly used to communicate and explain scientific concepts, yet their tendency to hallucinate poses signific...
By Raia Abu Ahmad, Nikolas Rauscher, Ekaterina Borisova, Fabio Barth, Georg Rehm, Sebastian M\"oller
arXiv:2608. 03860v1 Announce Type: cross Abstract: We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19.
By Kaysarul Anas Apurba, Md. Hasibul Hasan, Rofiqul Alam Shehab, Asab Azad
The paper investigates why retrieval‑based open‑ended evaluation fails in medical fact verification. By creating two detailed taxonomies—one for retrieval‑stage errors across five quality dimensions and another for verifier‑reasoning errors across six steps—the authors automatically label evidence quality and reasoning errors using an LLM‑as‑Judge pipeline. Their large‑scale stress tests across multiple retrieval methods and verifier models show that increasing model size, reasoning effort, source breadth, or medical fine‑tuning does not eliminate these failure modes, indicating fundamental limits of the retrieve‑then‑verify paradigm in open‑ended medical contexts.
By Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jim\'enez Guti\'errez, Mark Dredze
arXiv:2607. 01131v1 Announce Type: cross Abstract: Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation.
By Bingchen Zhao, Sara Beery, Oisin Mac Aodha
arXiv:2604. 13201v2 Announce Type: replace-cross Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging.
By Oliver Bentham, Vivek Srikumar
The paper introduces VERA-RL, a reinforcement‑learning framework for proactive scientific error verification in academic papers. It builds on a Reason–Verify–Scan workflow and presents VERA‑13K, a 12,900‑sample dataset with 4,300 matched reasoning chains covering six error categories across natural‑science domains. The authors also define fine‑grained rewards for reasoning completeness, evidence alignment, and error precision, and show that training Qwen3‑VL‑8B with VERA‑RL improves verifiable reasoning to levels comparable with flagship multimodal large language models.
By Rongjin Li, Yuanxin Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun
The paper identifies a specific issue in supervised fine‑tuning (SFT) of large language models called factual access failure, where models can recognize correct facts under constrained tests but fail to generate them in open‑ended settings. It demonstrates that SFT can cause both genuine wrong answers and expression‑level errors such as verbosity or formatting mismatches. To mitigate this, the authors propose Recall‑Anchored Distillation (RAD), a self‑distillation method that aligns the fine‑tuned model with the base model’s soft output distribution on unlabeled out‑of‑distribution text, thereby recovering lost factual recall without needing labeled data.
By Haodong Chen, Yadong Wang, Shengtao Wen, Dong Liang, Xiang Chen
arXiv:2603. 03322v2 Announce Type: replace-cross Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery.
By Chaoqun Yang, Xinyu Lin, Shulin Li, Wenjie Wang, Ruihan Guo, Fuli Feng, Tat-Seng Chua
arXiv:2605.22878v2 Announce Type: replace
Abstract: Artificial intelligence is rapidly entering the core workflows of scientific research. Yet reliable scientific reasoning requires access to accumul...
By Shuofei Qiao, Yunxiang Wei, Busheng Zhang, Mengru Wang, Jiazheng Fan, Huadong Jian, Bin Wu, Shumin Deng, Yida Xue, Zifan Cheng, Xiang Chen, Dan Zhang, Junfeng Fang, Ningyu Zhang, Keyan Ding, Qiang Zhang, Jeff Z. Pan, Emine Yilmaz, Huajun Chen