ReFIne is a training framework that augments large reasoning models with three trustworthiness properties: interpretability, faithfulness, and reliability. It combines supervised fine‑tuning with GRPO to produce structured, tag‑based reasoning traces, explicitly disclose decisive information, and provide self‑assessments of soundness and confidence. Applied to Qwen3 models, ReFIne improves interpretability by 44.0 %, faithfulness by 18.8 %, and reliability by 42.4 % on mathematical benchmarks.
By Chung-En Sun, Ge Yan, Akshay Kulkarni, Tsui-Wei Weng
arXiv:2604. 11996v2 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy?
By Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi
arXiv:2607. 17266v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing.
By Peiji Yu, Xin Chen, Tianxing Wu
arXiv:2606. 03969v1 Announce Type: cross Abstract: Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode.
By Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu, Arman Cohan
arXiv:2606. 21678v2 Announce Type: replace-cross Abstract: Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reasoning.
By Vatsal Ananthula, Adarsh Kumarappan
The paper "Strategic Self-Consistency" investigates how large language model providers might exploit the self‑consistency technique—generating multiple reasoning paths and selecting the majority answer—to overcharge users. The authors present a simple, efficient algorithm that strategically generates and reorders extra reasoning paths so that each appears necessary for the majority vote, thereby evading detection by auditors. Experiments on Llama, Qwen, and DeepSeek-R1 models across math, science, and QA benchmarks show that the added paths follow a heavy‑tailed distribution and that significant overcharging can persist even under stringent audits with a false‑positive rate below 0.1.
By Tori Qiu, Ander Artola Velasco, Manuel Gomez-Rodriguez