arXiv AI By Jin Liu, Steffen Thoma, Achim Rettinger

TriQua: Reconciling Granularity and Context in Factuality Evaluation

Read the original on arXiv AI →

arXiv:2608. 05228v1 Announce Type: new Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 28

ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

ElementCheck is a new framework for evaluating the factuality of long-form text that addresses limitations of the traditional decompose‑retrieve‑verify pipeline. Rather than breaking sentences into atomic sub‑claims, it extracts entity pairs linked by verifiable connections to form an element graph, using the graph’s topology to gauge sentence complexity. This allows simple sentences to be verified directly while complex ones undergo targeted element‑level refinement, and the authors introduce the FastFact‑Sent benchmark to support fine‑grained evaluation, demonstrating consistent improvements across multiple backbone models.

By Xinming Wang, Haoran Du, Yi Chen, Jian Xu, Hongming Yang, Han Hu, Yulong Chen, Cheng-Lin Liu, Xu-Yao Zhang
arXiv Computation and Language
Aug 28

AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification

AEScorer is an agentic evidence‑grounded framework designed for graded factuality verification, addressing the limitation of binary judgments in current methods. It operates in two stages: first, it gathers and refines external evidence through agentic search; second, it predicts a scalar factuality score to capture nuanced differences in correctness. The authors also introduce GradedVeriBench, a benchmark covering general and multi‑hop question answering, and demonstrate that AEScorer outperforms existing approaches on this new benchmark.

By Hui Huang, Muyun Yang, Yuki Arase
arXiv AI
Sep 10

Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation

The paper introduces Fact-Ablated Evaluation (FAE), a framework that iteratively removes cited evidence to test whether large language models (LLMs) adjust their fact‑checking predictions accordingly. Experiments reveal that many off‑the‑shelf LLMs rely more on internal knowledge than on the provided evidence. To address this, the authors propose REAL, a training method that uses counterfactual evidence supervision to encourage LLMs to base veracity judgments on evidence, achieving better evidence‑dependent performance across four datasets.

By Xingyu Deng, Mingzi Cao, Nikolaos Aletras, Xi Wang, Mark Stevenson