arXiv AI By Brian Rabern, Philipp Mondorf, Barbara Plank

LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models

Read the original on arXiv AI →

LogicSkills is a benchmark designed to isolate three core logical abilities in large language models: formal symbolization, countermodel construction, and validity assessment. The dataset draws items from the two-variable fragment of first‑order logic without identity, presented in both English and a Carrollian nonce‑word language, and all instances are solver‑verified with Z3. Results show that conventional instruction‑tuned LLMs excel at validity assessment but struggle with symbolization and countermodel construction, whereas recent reasoning‑tuned models perform well across all tasks, indicating a more systematic logical skill profile.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen