arXiv AI By Riccardo Revalor, Jalees Rehman, Debjit Pal

Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

Read the original on arXiv AI →

arXiv:2607. 08017v1 Announce Type: cross Abstract: Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to detect as they evaluate only final-answer agreement while ignoring the logical validity of intermediate steps.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability

ReFIne is a training framework that augments large reasoning models with three trustworthiness properties: interpretability, faithfulness, and reliability. It combines supervised fine‑tuning with GRPO to produce structured, tag‑based reasoning traces, explicitly disclose decisive information, and provide self‑assessments of soundness and confidence. Applied to Qwen3 models, ReFIne improves interpretability by 44.0 %, faithfulness by 18.8 %, and reliability by 42.4 % on mathematical benchmarks.

By Chung-En Sun, Ge Yan, Akshay Kulkarni, Tsui-Wei Weng
arXiv AI
6d ago

Strategic Self-Consistency

The paper "Strategic Self-Consistency" investigates how large language model providers might exploit the self‑consistency technique—generating multiple reasoning paths and selecting the majority answer—to overcharge users. The authors present a simple, efficient algorithm that strategically generates and reorders extra reasoning paths so that each appears necessary for the majority vote, thereby evading detection by auditors. Experiments on Llama, Qwen, and DeepSeek-R1 models across math, science, and QA benchmarks show that the added paths follow a heavy‑tailed distribution and that significant overcharging can persist even under stringent audits with a false‑positive rate below 0.1.

By Tori Qiu, Ander Artola Velasco, Manuel Gomez-Rodriguez