The paper examines the reliability of lie detection probes for language models when the models adopt anti-factual personas, such as conspiracy theorists. A dataset of 8,916 human-reviewed responses from three LLMs was created, and eight existing probes were evaluated, revealing many fail to flag falsehoods under these personas. The authors also constructed confounder datasets showing that probes often track spurious correlations like instruction compliance, and propose a simple linear probe that performs best on both persona and confounder tests.
By Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart B\"urger
The paper presents a theory-informed computational framework that converts cross-disciplinary theories of fake news into measurable features for automated detection and explanation. By reviewing theories from social sciences, psychology, economics, and more, the authors establish a broad theoretical foundation for computational modeling. Experiments on benchmark datasets demonstrate that theory-derived features are predictive, provide interpretable diagnostic signals, and that multi-feature models generally outperform individual features, though gains are modest.
By Zhaoyang Cao, Miriam Metzger, Reza Zafarani
arXiv:2609.15561v1 Announce Type: new
Abstract: Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy?...
By Sarra Gharsallah, Adele Robaldo, Mariia Tokareva, Giovanni Gatti Pinheiro, Ilyana Guendouz, Rapha\"el Troncy, Paolo Papotti, Pietro Michiardi
arXiv:2606. 26437v1 Announce Type: cross Abstract: Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist.
By Siyi Liu, Aaron Halfaker, Dan Roth, Patrick Xia
arXiv:2606. 18060v1 Announce Type: new Abstract: As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important.
By Xinyang Liao, Lingyu Li, Huacan Liu, Tianle Gu, Yang Yao, Tong Zhu, Yan Teng, Yingchun Wang
arXiv:2208. 11582v2 Announce Type: replace-cross Abstract: The wide spread of false information online, including misinformation and disinformation, has become a major problem for our highly digitised and globalised society.
By Haiyue Yuan, Enes Altuncu, Shujun Li, Can Baskent, Jason R. C. Nurse
The paper investigates whether trustworthiness scores and truth judgments produced by LLM-as-Judge systems are truly independent. Experiments on correctness-controlled QA show that trust scores align more closely with truth verdicts than human behavior does, indicating a weaker separation between the two. Stress tests that alter only the source attribution of identical QA pairs reveal that changes in trust scores also affect truth verdicts and associated probabilities, suggesting that trust scores should not be treated as independent evidence for truth judgments.
By Xin Sun, Di Wu, Yuchen Guo, Jiahuan Pei, Isao Echizen, Abdallah El Ali, Saku Sugawara
arXiv:2606. 12618v1 Announce Type: new Abstract: Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires testbeds where models verifiably believe the opposite of what they say.
By Alan Cooney, David Africa, Geoffrey Irving
arXiv:2609.24369v1 Announce Type: cross
Abstract: Deception plays a central role in Intelligence operations, yet it remains difficult to analyse systematically without expert knowledge of reasoning p...
By Stefan Sarkadi, Xabier Garmendia, Jack Mumford, Trevor Bench-Capon
arXiv:2608. 12852v1 Announce Type: cross Abstract: Language can describe states of affairs that are false and states of affairs that could not be the case at all.
By Yoon Pyo Lee
arXiv:2607. 14152v1 Announce Type: cross Abstract: The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models.
By Michael Correll, Lucy Havens, Mahsan Nourani
The article "Beyond RAGs: Building Actually Truthful AI Harnesses" discusses the limitations of Retrieval-Augmented Generation (RAG) systems, emphasizing that retrieval alone does not guarantee evidence for AI claims. It explores methods for constructing AI systems that can substantiate their statements, moving beyond simple retrieval to more robust proof mechanisms. The piece highlights the importance of developing AI that can verify its own outputs rather than merely retrieve information.
By Ari Joury, PhD