The paper investigates whether large language models (LLMs) follow Occam's Razor when performing inductive and abductive reasoning. It introduces a synthetic framework for generating questions that require both types of reasoning and a new automated metric to evaluate the simplicity and correctness of generated hypotheses. Experiments show that while LLMs can handle simple scenarios, they struggle with complex world models and producing high‑quality, simplest hypotheses, even when using advanced reasoning techniques.
By Yunxin Sun, Abulhair Saparov
arXiv:2609.13808v1 Announce Type: new
Abstract: Structured knowledge fact checking aims to determine the truthfulness of natural language claims by reasoning over structured evidence. Recent program-...
By Yifei Li, Xiaohan Zheng, Wentao Qian, Liansheng Zhuang
The paper introduces Multi-Agent Agentic Graph Learning (MAAGL), a framework that partitions a graph into communities and assigns a dedicated agent to each community for specialized reasoning. MAAGL addresses two key challenges in existing agentic graph learning: it preserves permutation invariance by summarizing structural evidence with a dynamic structural signature, and it controls context size by filtering semantic evidence to the top‑k relevant nodes. Experiments on four benchmark datasets demonstrate that MAAGL outperforms state‑of‑the‑art agentic graph learning methods.
By Liang Qu, Jianxin Li, Hua Wang
arXiv:2606. 04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks.
By Leonardo Bertolazzi, Katya Tentori, Raffaella Bernardi
arXiv:2607. 14149v1 Announce Type: new Abstract: Although large language models (LLMs) have set benchmarks for zero-shot reasoning, their deployment remains cost-prohibitive and environmentally taxing.
By Dimitrios Kelesis, Konstantinos Bougiatiotis, Georgios Paliouras
GraphCert introduces a method to bootstrap graph reasoning agents by generating graph‑grounded question‑answer pairs and certifying the supporting evidence. The approach uses a Bootstrapped Graph Quizzer to produce QA pairs, then executes and semantically curates the evidence into certified rubrics that guide reward‑based training of a Graph Solver. Experiments on five GRBENCH domains show GraphCert outperforms larger LLM agents and demonstrates robust policy transfer across heterogeneous graphs.
By Weiqi Jiang, Yuchen Ying, Rui Wang, Kaixuan Chen, Bingde Hu, Shunyu Liu, Yu Wang, Tongya Zheng
arXiv:2604.27251v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicite...
By Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Yuxiang Zhou, Maria Liakata, Nikolaos Aletras
The paper argues that pattern recognition and step‑by‑step reasoning lie on a spectrum, with large language models (LLMs) learning the latter when the next token depends on only a few preceding tokens. It formalises reasoning traces as paths on a De Bruijn graph, showing that the number of edges is far smaller than the number of possible traces, making step‑by‑step reasoning sample‑efficient. Experiments fine‑tuning Qwen2.5‑1.5B‑Instruct demonstrate that a moderate density of states balances accuracy and robustness, and that real‑world models like Qwen3 retain most of their performance even when attention is limited to a small sliding window.
By Amrut Nadgir, Pratik Chaudhari, Vijay Balasubramanian
arXiv:2606. 13220v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving.
By Fabrizio Marozzo, Pietro Li\`o
arXiv:2606. 16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit.
By Jiajie Jin, Zhao Yang, Wenle Liao, Yuyang Hu, Guanting Dong, Xiaoxi Li, Yutao Zhu, Zhicheng Dou
arXiv:2607. 23019v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound.
By Zirong Chen, Meiyi Ma
The paper introduces Procedural Graphs, a framework that structures procedural knowledge for large language model agents as (procedure, relation, procedure) triplets, analogous to knowledge graphs for factual data. At each decision point, a guidance model uses the local subgraph to bias the agent’s next action, while an LLM refiner self‑evolves the graph by comparing failed and successful trajectories, editing its topology to improve performance. Experiments across various datasets, tasks, and LLMs show that Procedural Graphs consistently outperform memory‑based baselines, and the self‑evolution mechanism further enhances results without manual engineering.
By Yuxing Lu, Yicheng Chen, Shanchan Wu, Sercan \"{O}. Ar{\i}k