Making Implicit Premises Explicit in Logical Understanding of Enthymemes
arXiv:2603. 06114v2 Announce Type: replace-cross Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.
The paper presents a neuro‑symbolic approach for selecting the correct omitted component in enthymemes, extending prior work from missing‑premise to missing‑claim selection. It replaces binary entailment with logical‑resistance scores and introduces the Possible‑World Atom‑Link Formalization (PWAL), which marginalizes over alternative semantic‑link configurations while keeping translated formulae fixed. Experiments on five tasks show that PWAL improves strict accuracy by up to 30.86 percentage points and reduces tie rates significantly, while also providing a transparent trace of each comparison.
arXiv:2603. 06114v2 Announce Type: replace-cross Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.
arXiv:2606. 03969v1 Announce Type: cross Abstract: Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode.
arXiv:2607. 14349v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility.
arXiv:2605. 04539v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference signals from human annotators or LLM judges exhibit a systematic verbosity bias that rewards fluency over logical correctness.
arXiv:2507. 09751v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but exhibit problems with logical consistency in their output.
arXiv:2510. 00492v3 Announce Type: replace Abstract: The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic.
arXiv:2607. 28680v1 Announce Type: cross Abstract: Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities.
arXiv:2608. 08514v1 Announce Type: new Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models).
The paper tackles the challenge of reconstructing argument graphs from natural language by addressing enthymemes—arguments with implicit premises. It proposes a neuro‑symbolic pipeline that employs large language models to generate intermediate implicit premises, translates them into logical formulas, and combines them with explicit premises and claims to determine entailment, contradiction, or neutrality. The method is evaluated on the Microtext Argumentative Corpus.
arXiv:2606. 02837v1 Announce Type: cross Abstract: Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), making the quality of NL-to-FOL benchmarks essential -- yet these datasets have never been rigorously audited.
arXiv:2607. 11266v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs substantial computational costs due to over-reasoning: the generation of redundant, verbose, or irrelevant steps.
arXiv:2606. 16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit.