arXiv:2507. 09751v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but exhibit problems with logical consistency in their output.
By Bradley P. Allen, Prateek Chhikara, Thomas Macaulay Ferguson, Filip Ilievski, Paul Groth
The paper introduces OBJECTION, an inference-time pipeline that adds an Adversarial Lawyer Agent to each of the three reasoning steps—offense, unlawfulness, and culpability—in legal judgment prediction models. By actively injecting defense arguments, the agent challenges the model’s default assumption of guilt, which is common in datasets biased toward guilty outcomes. Using a new Natural Innocent dataset of 3.4k real cases, OBJECTION reduces the False Guilty Rate from 82.93% to 16.69%, demonstrating significant improvement in substantive legal reasoning.
By Jaehoon Jeong, Jay-Yoon Lee
arXiv:2607. 04096v1 Announce Type: new Abstract: Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window.
By Vishvesh Bhat, Jay Vaghasiya, Emmanuel Anaya Gonzalez
arXiv:2606. 09030v1 Announce Type: cross Abstract: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk scores for patient triage and interpretable rationales that clinicians can verify.
By Hyeongwon Jang, Gyouk Chu, Changhun Kim, Joonhyung Park, Hangyul Yoon, Eunho Yang
arXiv:2608.22483v1 Announce Type: new
Abstract: Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is mis...
By Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Li Chen
The paper introduces contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), a formalism used to represent and reason with information. Unlike traditional explanations that focus on a single argument, contrastive explanations highlight the differences between two topic arguments. The authors propose a general form of contrastive attribution functions (CAFs), present CAFs based on removal, gradients, and Shapley-values, and demonstrate their applicability in healthcare and bias identification contexts.
By Xiang Yin, Nico Potyka, Antonio Rago, Francesca Toni
arXiv:2606. 15646v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed natural language processing, but their lack of interpretable reasoning and tendency to hallucinate pose significant challenges for legal applications.
By Deepa Tilwani, Yash Saxena, Ankur Padia, Srinivasan Parthasarathy, Manas Gaur
arXiv:2608.29529v1 Announce Type: cross
Abstract: Semantic alignment between specialized normative texts is challenging when equivalent requirements use different terms, syntax, and levels of abstrac...
By William Schroeder
The paper introduces a framework that uses large language models (LLMs) to generate natural‑language narratives explaining cross‑sectional stock return predictions. It combines temporal Shapley additive explanations (SHAP) from an XGBoost model with historical regime analogs to provide context. A controlled study shows that progressively externalizing numerical and relational reasoning improves evidence faithfulness and accuracy, while historical analogs boost human‑rated usefulness.
By Sujung Kim, Seung Hwan Cho, Sangjin Park, Young-Min Kim
arXiv:2606. 24414v1 Announce Type: new Abstract: Formal verification produces machine-checkable certificates that attest to the satisfaction or violation of temporal properties, yet these certificates remain opaque to non-specialist stakeholders.
By Andoni Rodriguez, Alberto Pozanco, Daniel Borrajo
Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window. The prevailing route to improve such reasoning is test-time scaling, which trains models to search over long chains of thought; but the resulting capability is entangled in model weights, is not verifiable step-by-step, and is costly at inference.
arXiv:2609.24480v1 Announce Type: cross
Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the conv...
By Kalash Shah, Kunal Singh, Snehan J, Shreyas Singh