arXiv:2608. 18242v1 Announce Type: new Abstract: We introduce ClosureBench, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth.
By Stefano Goria (AIM Research Lab)
arXiv:2608. 08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said.
By Fengrong Wan, Chengcan Wu, Ningtao Lyu
arXiv:2608.28978v1 Announce Type: new
Abstract: Knowledge graphs have been proposed as a structured alternative to flat retrieval-augmented generation for long-term agent memory, on the assumption th...
By Theo Rusu, Sourena Khanzadeh, Manar Alalfi
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2609.38340v1 Announce Type: new
Abstract: When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accu...
By Jiajun Wu, Jian Yang, Zixiang Ni, Zhenzhu Li, Bin Chong
arXiv:2606.22419v3 Announce Type: replace
Abstract: A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical benc...
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2606. 30128v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning, but the source is contested: do the intermediate steps help because they carry useful semantic content, or because conditioning on more tokens buys extra computation before the model commits to an answer?
By Wenlong Wang, Fergal Reid
The paper evaluates two large language models, Claude Sonnet 4.5 and Claude Opus 5, on the bidirectional English Resource Grammar (ERG) tasks of generating English from Minimal Recursion Semantics (MRS) and parsing English into MRS. In generation, Opus achieves 76.3 BLEU—surpassing a 72k‑pair trained system and matching a million‑pair system—while Sonnet scores 65.7 BLEU, rising to 69.6 when selecting from ACE’s candidates. In parsing, both models lag behind ACE, attaining only 57.2 and 65.5 F₁ respectively, with exact‑match on about 1 % of sentences, highlighting that high generation scores do not guarantee accurate semantic parsing.
By Soham Dan
The paper investigates how small language models (SLMs) perform in knowledge graph question answering (KGQA) when evaluated on the reasoning paths they take, rather than just the final answer. Using the THESEUS navigation and traceability framework, the authors test frozen, off‑the‑shelf SLMs as local action policies that choose graph actions and decide when to stop, without any task‑specific training or free‑form answer generation. By measuring both Hits@1 and Path Edit Distance (PED) across the Kinship and MQuAKE‑ST datasets, the study finds that models vary significantly in both answer accuracy and path fidelity, and that prompting can either help or hurt navigation depending on the model.
"whyItMatters":"The results show that evaluating SLMs solely on endpoint accuracy can be misleading, highlighting the need to assess reasoning path fidelity in KGQA tasks."
By Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini
arXiv:2608. 14808v1 Announce Type: new Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that information determines a unique answer.
By Yepeng Huang, Jiawen Zhang, Michelle Dai, Xiaorui Su, Shanghua Gao, Zi Wang, Marinka Zitnik
arXiv:2606. 08471v1 Announce Type: cross Abstract: Recently, language models have made rapid progress across various domains and applications.
By Marina Igitkhanian, Erik Arakelyan
The paper investigates multi‑hop question answering systems and identifies two distinct failure modes: retrieval failures, where the necessary passage is not retrieved, and extraction failures, where the passage is retrieved but the required fact cannot be extracted—a phenomenon termed the fact‑grounding gap. Across three standard benchmarks, extraction failures account for nearly half of all per‑hop deficiencies and are invisible to standard retrieval metrics, remaining unresolved by retrieval‑only interventions. The study shows that these two bottlenecks require different solutions, a distinction currently missing from evaluation practices.
By Kevin Mo, Nathan Mo, Richard Zhu