arXiv AI

Human and AI-generated texts between modal logic and statistics

The paper examines the geometry of semantic neighbourhood graphs as modal logic, translating this view into a statistical framework to differentiate human from AI-generated text. It models texts as worlds in a finite frame where accessibility is defined by the k‑nearest‑neighbour relation of transformer embeddings, and measures the frequencies of modal axioms (B, 4, 5, D) as validation degrees. A prompt‑balanced comparison shows consistently higher degrees for axioms 4 and 5 in AI‑generated corpora, and the study further introduces measures of groundedness and situatedness, recasting the analysis in a sequent‑style tableau setting. "whyItMatters":"The work provides a quantitative, modal‑logic‑based method to detect structural differences between human and machine‑generated text, offering a new lens for evaluating AI language models."

arXiv Computation and Language
Sep 22

A vector logic for intensional formal semantics

arXiv:2602.02940v2 Announce Type: replace-cross Abstract: Formal semantics and distributional semantics are distinct approaches to linguistic meaning: the former models meaning as reference via model...

By Daniel Quigley
arXiv AI
Jun 2

KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models

arXiv:2604. 17621v2 Announce Type: replace Abstract: Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg.

By Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang, Johan Bos
arXiv Computation and Language
Sep 22

Euston: Training Away Mathematical Sycophancy Without Losing the Mathematics

Euston is an 8‑B parameter mathematical claim‑verification model that resists producing false derivations when presented with corrupted theorems. It was trained on 3,026 matched true/corrupted statement pairs generated by GraphSynth, a probabilistic factor‑graph generator, and fine‑tuned from DeepSeek‑R1‑8B using GRPO. On a balanced held‑out split, Euston’s balanced accuracy rose from 29.50 % to 63.75 %, and its discrimination gap improved from –0.5 % to +27.5 %, while maintaining comparable general mathematical ability and reducing response length and truncation rates.

By Zehua Cheng, Wei Dai, Jiahao Sun