arXiv AI By Brett Reynolds

Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models

Read the original on arXiv AI →

arXiv:2608. 14252v1 Announce Type: new Abstract: Recent work suggests that some large language model representations have content or reference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
5d ago

Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation

The paper introduces TRACE, a fine‑tuning framework for Retrieval‑Augmented Generation (RAG) that addresses conflicts between retrieved knowledge and a model’s internal knowledge. TRACE uses multi‑agent debate traces to identify correct and incorrect candidates and answer‑shift patterns, providing fine‑grained supervision for reliable knowledge‑source selection. It also incorporates an answer‑completeness regularization mechanism to prevent empty, overly short, or prematurely terminated responses, thereby improving robustness against misleading retrieved content and enhancing answer quality.

By Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu, Ji Chen, Xing Wang
arXiv AI
Sep 25

PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.

By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv Machine Learning
Sep 21

Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

arXiv:2609.20973v1 Announce Type: cross Abstract: Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-ste...

By Jiazhang Cai, Tao Wang, Ruidong Zhang, Siyuan Li, Terry Ma, Luyang Fang, Haoran Lu, Huimin Cheng, Yingchuan Zhang, Shushan Wu, Rui Xie, Lin Tang, Chao Huang, Rongjie Liu, Ziyu Liu, Meizhi Yu, Yongkai Chen, Yifan Zhou, Zeliang Sun, Chang Liu, Zhen Xiang, Wei Xiao, Zixin Rao, Xinyi Liu, Yutong Hu, Mengrui Zhang, Jing Zhang, Weidi Luo, Jincheng Yu, Zhengliang Liu, Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen, Xinliang Li, Tianming Liu, Wenxuan Zhong, Ping Ma