arXiv Computation and Language By Svetlana Churina, Kokil Jaidka, Anab Maulana Barik, Harshit Aneja, Cai Yang, Insyirah Binte Imam Mujtahid, Wynne Hsu, Mong Li Lee

Althea: The Fact-Checking--Metalearning Tradeoff in AI-Assisted Verification

Read the original on arXiv Computation and Language →

Althea is a retrieval‑augmented system that supports user‑driven claim evaluation, matching standard pipelines on AVeriTeC while improving discrimination between supported and refuted claims. In a longitudinal survey experiment with 961 participants, two AI‑assisted treatments—Exploratory (guided reasoning) and Summary (synthesized verdicts)—initially boosted accuracy and confidence, but these gains faded after the system was removed, leaving no advantage over unrelated news. In contrast, a Self‑search baseline, which lacks a fading procedure, maintained a significant advantage, highlighting a fact‑checking–metalearning tradeoff where methods that improve immediate accuracy may not foster durable literacy gains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 22

EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy

arXiv:2609.22223v1 Announce Type: cross Abstract: Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing c...

By Kening Zheng, Aoying Zheng, Zhigang Chang, Yazhi Guo, Miaotian Guo, Qingwei Zong, Xianhai Xie, Weiqiang Jin, Chengze Li, Hanrong Zhang, Jie Yang, Wei-Chieh Huang, Lingzhe Zhang, Liancheng Fang, Xin Zou, Hanqian Li, Jiahao Huo, Yibo Yan, Zizhuang Deng, Lei Miao, Wei Guo, Haihong Tang, Bo Zheng, Philip S. Yu
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen