arXiv AI

REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

REPAIR is a self‑evolving data augmentation framework designed to improve scientific dense retrievers by addressing long‑tail concept gaps and fact sensitivity. It iteratively generates training data through diagnosis of long‑tail concepts, API‑guided evidence expansion, and hard negative mining, thereby grounding retrieval in factual reality. Experiments show that REPAIR outperforms 19 strong baselines across nine materials science and biomedical benchmarks.

arXiv AI
Aug 24

Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning

The paper introduces the task of Scientific Claim Unlearning and presents a new benchmark, SciUnlearn, to evaluate it. It highlights that language models trained on static scientific corpora risk disseminating outdated or retracted claims as scientific knowledge evolves. Current machine unlearning methods fail to effectively remove claim-level knowledge, often only suppressing it superficially, underscoring the need for specialized techniques for structured knowledge removal.

By Snigdha Paul, Manasi Patwardhan, Arman Cohan
arXiv Computation and Language
6d ago

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

The paper investigates why retrieval‑based open‑ended evaluation fails in medical fact verification. By creating two detailed taxonomies—one for retrieval‑stage errors across five quality dimensions and another for verifier‑reasoning errors across six steps—the authors automatically label evidence quality and reasoning errors using an LLM‑as‑Judge pipeline. Their large‑scale stress tests across multiple retrieval methods and verifier models show that increasing model size, reasoning effort, source breadth, or medical fine‑tuning does not eliminate these failure modes, indicating fundamental limits of the retrieve‑then‑verify paradigm in open‑ended medical contexts.

By Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jim\'enez Guti\'errez, Mark Dredze
arXiv Computation and Language
Aug 28

Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper

The paper introduces VERA-RL, a reinforcement‑learning framework for proactive scientific error verification in academic papers. It builds on a Reason–Verify–Scan workflow and presents VERA‑13K, a 12,900‑sample dataset with 4,300 matched reasoning chains covering six error categories across natural‑science domains. The authors also define fine‑grained rewards for reasoning completeness, evidence alignment, and error precision, and show that training Qwen3‑VL‑8B with VERA‑RL improves verifiable reasoning to levels comparable with flagship multimodal large language models.

By Rongjin Li, Yuanxin Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun
arXiv Machine Learning
Sep 3

When Literature Data Mislead Artificial Intelligence in Materials Discovery

The article examines how scientific literature, often used as a data source for AI in materials science, can contain hidden inaccuracies such as text-figure mismatches, ambiguous axis labels, unit inconsistencies, and missing measurement context. By tracing solid electrolyte conductivity values from original papers to curated datasets, the authors uncover recurrent errors that are numerically plausible yet hard to detect, leading to significant label noise in AI models. A cross-database example demonstrates that ambiguous reporting can cause a 100‑fold error in conductivity values, underscoring the need for traceable reporting, rigorous curation, and validation practices in AI-driven discovery.

By Qian Wang, Ying Li, Ryuhei Sato, Hidemi Kato, Shin-ichi Orimo, Hao Li, Eric Jianfeng Cheng