arXiv AI
Sep 24

Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

The paper investigates how large language models (LLMs) interpret negation across different logical semantics—open‑world vs. closed‑world, two‑ vs. three‑valued, and credulous vs. skeptical reasoning. Using the newly introduced NAFBench, a procedural generator that creates solver‑certified logic programs and their natural‑language verbalizations, the authors evaluate LLMs on four semantic viewpoints (SLDNF, well‑founded semantics, and stable‑model semantics). Results show a persistent gap: even the strongest models achieve only 59–74% accuracy, with many models sensitive to rule ordering and prone to overcommitment on undefined cases, though some frontier models reach near‑perfect performance on a fixed‑complexity set. "whyItMatters":"The study highlights that current LLMs struggle to reliably follow explicitly specified negation semantics, underscoring a limitation in their logical reasoning capabilities."

By Qiming Bao, Agnieszka Mensfelt, Michael J. Witbrock, Kostas Stathis
arXiv Computation and Language
Sep 16

Autoformalizing Argumentative Material Inferences

The paper introduces GUARD, a neuro‑symbolic system that autoformalizes argumentative material by completing missing premises (guards) before formal verification. It uses large language models to generate candidate guards, Isabelle/HOL to verify them, and a contrastive test to ensure the proof depends on the original premises and does not over‑generalize. Experiments on Debatepedia and ARCT show that GUARD improves verified‑faithful scores by over 30 points and reduces leakage by about 20 points compared to prior LLM‑driven theorem proving methods.

By Xin Quan, Reto Gubelmann, Andr\'e Freitas