arXiv AI By Siyi Liu, Aaron Halfaker, Dan Roth, Patrick Xia

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence

Read the original on arXiv AI →

arXiv:2606. 26437v1 Announce Type: cross Abstract: Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.