The paper audits the widely used ISOT/Kaggle Fake and Real News corpus and finds that extremely high reported accuracies (≈0.98) are largely due to shortcut signals rather than genuine veracity detection. A simple TF‑IDF linear classifier achieves perfect F1 when using only subject metadata, and even after removing metadata, newswire tags, and duplicate documents, the F1 drops only modestly, indicating that editorial style rather than specific tokens drives performance. Under topic‑disjoint and temporal transfer tests, performance collapses, and models transfer poorly to the independent LIAR benchmark, showing that within‑corpus scores reflect source and topic separability, not truth verification.
whyItMatters:"The study demonstrates that current high accuracy metrics on this fake‑news dataset are misleading, highlighting the need for more robust evaluation protocols that guard against shortcut learning."
By Yuvraj Verma
arXiv:2606. 19819v1 Announce Type: cross Abstract: Decomposing compound sentences into atomic, verifiable claims is a prerequisite for reliable automated fact-checking.
By Phuong Huu Vu Tran, Thuan Duc Mai, Bach Xuan Le
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
By Nicol\'as Vera Z\'u\~niga
arXiv:2609.15369v1 Announce Type: new
Abstract: Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-l...
By Jochen Madler (Sitefire)
arXiv:2609.09696v1 Announce Type: new
Abstract: Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poor...
By Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar, Medina Maloku
TRACE is a system designed to bridge the grounding contract gap in LitTraceQA by combining target-aware retrieval, independent typed evidence localization, multimodal table extraction, and schema-driven table construction. It indexes 27,487 papers using multiple representations while preserving question targets, predicts observation units for tables, and assembles rows with evaluator-compatible key normalization. On the official test set, TRACE achieves a 0.760613 overall score, with high paper F1, evidence F1, and multiple-choice accuracy, though table-row and macro cell performance remain lower.
By Sachin Gupta, Divya Godara
arXiv:2608. 00144v1 Announce Type: new Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines separate members from non-members from surface text alone.
By Victor Maricato
The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.
By Qing Ye, Meng-Hsuan Lin
arXiv:2608.21230v1 Announce Type: cross
Abstract: Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future sessions that match it. We measu...
By Arulnidhi Karunanidhi
arXiv:2608. 07946v1 Announce Type: cross Abstract: Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean.
By Mike Helwig
arXiv:2606. 23989v1 Announce Type: cross Abstract: End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify.
By Shuo Guan
arXiv:2609.24885v1 Announce Type: new
Abstract: When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved...
By John J. O'Hare