arXiv AI

What Was Said, Not What Was 'Thought': Type-6 Logic for CoT Verification

arXiv AI
Sep 24

Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

The paper investigates how large language models (LLMs) interpret negation across different logical semantics—open‑world vs. closed‑world, two‑ vs. three‑valued, and credulous vs. skeptical reasoning. Using the newly introduced NAFBench, a procedural generator that creates solver‑certified logic programs and their natural‑language verbalizations, the authors evaluate LLMs on four semantic viewpoints (SLDNF, well‑founded semantics, and stable‑model semantics). Results show a persistent gap: even the strongest models achieve only 59–74% accuracy, with many models sensitive to rule ordering and prone to overcommitment on undefined cases, though some frontier models reach near‑perfect performance on a fixed‑complexity set. "whyItMatters":"The study highlights that current LLMs struggle to reliably follow explicitly specified negation semantics, underscoring a limitation in their logical reasoning capabilities."

By Qiming Bao, Agnieszka Mensfelt, Michael J. Witbrock, Kostas Stathis
arXiv Computation and Language
Sep 16

Autoformalizing Argumentative Material Inferences

The paper introduces GUARD, a neuro‑symbolic system that autoformalizes argumentative material by completing missing premises (guards) before formal verification. It uses large language models to generate candidate guards, Isabelle/HOL to verify them, and a contrastive test to ensure the proof depends on the original premises and does not over‑generalize. Experiments on Debatepedia and ARCT show that GUARD improves verified‑faithful scores by over 30 points and reduces leakage by about 20 points compared to prior LLM‑driven theorem proving methods.

By Xin Quan, Reto Gubelmann, Andr\'e Freitas
arXiv AI
Aug 25

Walking on the DARKSIDE

arXiv:2608.23370v1 Announce Type: new Abstract: Large Language Models (LLMs) recognise patterns but do not natively track the path of exclusions that a coherent discourse demands. When an input rests...

By Aldo Gangemi, Emanuele Bottazzi
arXiv AI
Aug 11

Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing

arXiv:2608. 08514v1 Announce Type: new Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models).

By Minhan Cho, Jimin Kweon