arXiv AI By Hira Beril Kucuk, Norman W Paton, Jiaoyan Chen, Zhenyu Wu

Single and Multi Truth Data Fusion using Large Language Models

Read the original on arXiv AI →

arXiv:2606. 28062v1 Announce Type: cross Abstract: Data fusion, also known as truth discovery, is a data integration problem that aims to determine the correct value or set of values for each attribute of an object when presented with potentially conflicting values from multiple sources.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

ARAFA: An LLM-Generated Arabic Fact-Checking Dataset

A new large-scale Arabic fact‑checking dataset called Arafa has been created using an automated pipeline that generates claims from Arabic Wikipedia, mutates them into counterfactuals, and validates them against supporting or refuting evidence. The dataset contains 181,976 claim‑evidence pairs labeled as supported, refuted, or not enough information, and human evaluation shows high inter‑annotator agreement and strong validation accuracy. Fine‑tuned transformer models on Arafa achieve a Macro F1‑score of 77%, demonstrating its usefulness for Arabic fact‑checking tasks.

By Christophe Khalil, Shady Elbassuoni, Rida Assaf