arXiv AI By Daeyong Kwon, Soyoung Yoon, Seung-won Hwang

SAFE: An LLM-as-Verifier Framework for Evidence-Grounded Multi-Hop Reasoning

Read the original on arXiv AI →

arXiv:2604. 01993v2 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through invalid intermediate reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

arXiv:2603.16654v3 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, espe...

By Xiaojie Gu, Sherry T. Tong, Aosong Feng, Sophia Simeng Han, Jinghui Lu, Yingjian Chen, Yusuke Iwasawa, Yutaka Matsuo, Chanjun Park, Rex Ying, Irene Li