arXiv AI By Ting Wang, Yuanjie Shi, Yan Yan, Huan Zhang

Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models

Read the original on arXiv AI →

arXiv:2606. 08831v1 Announce Type: new Abstract: Large language models (LLMs) increasingly perform multi-step reasoning, where intermediate claims form implicit directed acyclic graphs whose node correctness is structurally conditioned on their ancestors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity

The paper introduces GUT, a method that uses directed acyclic graphs to represent all possible reasoning branches of Large Language Models (LLMs). It comprises two modules: GUT-Q, which quantifies reasoning uncertainty by approximating graph complexity, and GUT-O, which reduces uncertainty through reinforcement learning that rewards lower uncertainty. Experiments on four LLMs across five datasets demonstrate GUT’s effectiveness in measuring and mitigating reasoning uncertainty.

By Shuang Liang, Xin-Yu Hu, Xiang-Jun Ou, Shao-Qun Zhang
arXiv AI
Jun 24

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

arXiv:2605. 01482v3 Announce Type: replace Abstract: Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models , which may suffer from hallucinations and fractured logical chains.

By Yunhan Bu, Quan Zhang, Huaping Zhang, Guotong Geng, Chunxiao Gao, Askar Hamdulla, Juan Wang, Qiuchi Li, Baohua Zhang, Shuai Lei, Yunbo Cao, Zhunchen Luo
arXiv AI
Jun 16

VeriGraph: Towards Verifiable Data-Analytic Agents

arXiv:2606. 16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit.

By Jiajie Jin, Zhao Yang, Wenle Liao, Yuyang Hu, Guanting Dong, Xiaoxi Li, Yutao Zhu, Zhicheng Dou