arXiv AI By Anzhe Xie, Weihang Su, Jiaxin Mao, Yiqun Liu, Shaoping Ma, Qingyao Ai

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Read the original on arXiv AI →

arXiv:2606. 24894v2 Announce Type: replace-cross Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

arXiv:2606. 08036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence.

By Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma, Hao Tian, Siqi Zhou, Wenjing Gong, Kaili Zhang, Bingqian Chen, Mitch Zhang, Yifan Yang
arXiv AI
3d ago

Re-ranking and Late Interaction Drive Retrieval Quality: A Controlled Comparison of RAG Strategies for Scientific Question Answering

The paper presents a controlled comparison of six retrieval-augmented generation (RAG) strategies for scientific question answering on a large arXiv corpus. All pipelines use the same LLM generator and evaluation protocol, differing only in retrieval design—ranging from classic dense retrieval to late‑interaction methods like ColBERTv2. The authors also release a synthetic question dataset and code to enable reproducible, large‑scale evaluation of RAG trade‑offs.

By Bhagyesh Rathi, Eshan Chawla, William B. Andreopoulos