The paper introduces Fact-Ablated Evaluation (FAE), a framework that iteratively removes cited evidence to test whether large language models (LLMs) adjust their fact‑checking predictions accordingly. Experiments reveal that many off‑the‑shelf LLMs rely more on internal knowledge than on the provided evidence. To address this, the authors propose REAL, a training method that uses counterfactual evidence supervision to encourage LLMs to base veracity judgments on evidence, achieving better evidence‑dependent performance across four datasets.
By Xingyu Deng, Mingzi Cao, Nikolaos Aletras, Xi Wang, Mark Stevenson
arXiv:2608. 14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications.
By Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz
arXiv:2607. 11266v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs substantial computational costs due to over-reasoning: the generation of redundant, verbose, or irrelevant steps.
By Daeyeop Lee, Hwanjo Yu
arXiv:2608.30413v1 Announce Type: new
Abstract: Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of n...
By Jayanta Sadhu, Sayem Shahad, Kenneth Marino
arXiv:2607. 27405v2 Announce Type: replace-cross Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law.
By Nayera Hasan, Jack Greff, Alvin Grissom II
arXiv:2609.15145v1 Announce Type: new
Abstract: The reasoning ability of large language models (LLMs) is a critical factor for practical LLM-based applications. To investigate the current reasoning c...
By Runa Yoshida, Kosuke Nishida, Kyosuke Nishida
Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have exami...
arXiv:2601. 22642v2 Announce Type: replace Abstract: Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid.
By Chuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan, Zijian Zhao, Zhengyu Chen, Yuchen Tian, Lijun Wu, Conghui He, Sirui Han, Yike Guo
arXiv:2508. 10971v2 Announce Type: replace-cross Abstract: Knowledge graphs (KGs) can be enhanced through rule mining; however, the resulting logical rules are often difficult for humans to interpret due to their inherent complexity and the idiosyncratic labeling conventions of individual KGs.
By Nasim Shirvani-Mahdavi, Chengkai Li
The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially in inferential categories like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.
The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially on inferential tasks like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.
By Yutong Yao, Yanjie Cao, Guanhua Chen, Xu Yang, Junchao Wu, Zeyu Wu, Lidia S. Chao, Derek F. Wong
The paper introduces an adaptive triggering mechanism for bias correction in large language model (LLM) reasoning. By framing bias intervention as an online change‑point detection problem, the authors update a CUSUM statistic at each step using either a white‑box next‑token probability signal or a black‑box LLM judge signal, and inject corrective prompts only when the accumulated evidence exceeds a calibrated threshold. Experiments on gpt‑4o‑mini and six open‑weight models show that adaptive black‑box triggering restores most of the accuracy lost by fixed‑interval interventions while reducing the number of corrections, whereas the white‑box signal improves ambiguous‑item accuracy but can hurt disambiguated‑item accuracy due to difficulty distinguishing stereotype reliance from correct evidence.
By Nayoung Kim, Mickey Mancenido, Huan Liu