arXiv AI

THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA

arXiv:2607. 20459v1 Announce Type: cross Abstract: Multi-hop question answering requires retrieving and integrating evidence from multiple contexts.

arXiv Computation and Language
Sep 22

SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine

arXiv:2410.17021v2 Announce Type: replace Abstract: Large Language Models with chain-of-thought prompting, such as OpenAI-o1, have shown impressive capabilities in natural language inference tasks. H...

By Xiaochen Wang, Liang Chen, Reza Haf Zhe Yang, Yiru Wang, Xiangdi Meng, Kunhao Pan, Zhifang Sui, Junqing He
arXiv AI
Oct 1

Hierarchical Reasoning Model

arXiv:2506.21734v4 Announce Type: replace Abstract: Reasoning, the process of devising and executing complex goal-oriented action sequences, remains a critical challenge in AI. Current large language...

By Guan Wang, Jin Li, Yuhao Sun, Xing Chen, Changling Liu, Yue Wu, Meng Lu, Sen Song, Yasin Abbasi Yadkori
arXiv AI
Sep 30

Rethinking Reasoning Paths as Phase-Structured Trajectories

The paper proposes PAIR, a method that treats reasoning paths of large language models as phase‑structured trajectories within each question. By sampling multiple trajectories per question, aligning them to shared relative phases, and comparing successful versus unsuccessful paths only within the same phase, PAIR isolates path‑quality signals from question‑level variation. Experiments show that standard correctness probes lose predictive power under this within‑question evaluation, while PAIR improves trajectory ranking, Best‑of‑N selection, and enables phase‑wise steering of generation outcomes.

By Zhenghao He, Guangzhi Xiong, Sanchit Sinha, Bohan Liu, Wenqian Ye, Aidong Zhang
arXiv AI
Sep 24

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

The paper investigates how small language models (SLMs) perform in knowledge graph question answering (KGQA) when evaluated on the reasoning paths they take, rather than just the final answer. Using the THESEUS navigation and traceability framework, the authors test frozen, off‑the‑shelf SLMs as local action policies that choose graph actions and decide when to stop, without any task‑specific training or free‑form answer generation. By measuring both Hits@1 and Path Edit Distance (PED) across the Kinship and MQuAKE‑ST datasets, the study finds that models vary significantly in both answer accuracy and path fidelity, and that prompting can either help or hurt navigation depending on the model. "whyItMatters":"The results show that evaluating SLMs solely on endpoint accuracy can be misleading, highlighting the need to assess reasoning path fidelity in KGQA tasks."

By Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini
arXiv Computation and Language
Sep 21

MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

MIRAGE is a new inference-time framework that enhances large language models by using a Selector to choose effective conceptual perspectives and a Reasoner to solve tasks step-by-step, aggregating multiple perspectives when needed. It is inspired by human cognitive flexibility and is designed to improve performance on complex mathematical, scientific, and logical problems. Experiments on GSM8K, MATH500, MMLU-Pro, and Game-of-24 show that MIRAGE outperforms Chain-of-Thought and diverse prompting ensembles, boosting accuracy with minimal inference overhead.

By Arash Lagzian, Srinivas Anumasa, Dianbo Liu
arXiv AI
Sep 1

AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

AgenticRag‑R1 is a reinforcement‑learning framework that integrates reasoning, retrieval, and memory through a stack and fine‑grained action space. It uses hierarchical action‑aware rewards and an information‑aware trajectory rejection strategy to support long‑horizon learning. Experiments on multi‑hop, open‑domain, and agentic reasoning benchmarks show that AgenticRag‑R1 outperforms strong baselines and produces robust, interpretable, memory‑aware reasoning behaviors.

By Xinke Jiang, Yue Fang, Zhibang Yang, Jiaran Gao, Zhixin Zhang, Tao Feng, Rihong Qiu, Wentao Zhang, Hongxin Ding, Ruizhe Zhang, Yongxin Xu, Yuheng Huang, Xu Chu, Junfeng Zhao, Yasha Wang
arXiv AI
Aug 6

Chained Recursive Language Models for Multi-Iteration Reasoning

arXiv:2608. 05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer.

By Purbesh Mitra, Sennur Ulukus
arXiv Computation and Language
Aug 25

GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning

arXiv:2608.22479v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge-intensive questions. For complex multi-hop quest...

By Jun Chen, Yongchao Liu, Pengyu Qiu, Jiajun Zheng, Juelu Zhang, Yujie Zeng, Qin Zhang, Ziyue Qiao, Xiao Luo
arXiv Computation and Language
Aug 31

PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering

PRISM is an agentic retrieval framework that uses large language models in a structured loop to improve evidence gathering for multi‑hop question answering. It splits retrieval into three specialized agents—a Question Analyzer, a Selector focused on precision, and an Adder focused on recall—whose iterative interaction yields a compact yet comprehensive evidence set. Experiments on HotpotQA, 2WikiMultiHopQA, MuSiQue, and MultiHopRAG show that PRISM consistently outperforms strong baselines by achieving higher retrieval accuracy and filtering out distracting content.

By Md Mahadi Hasan Nahid, Davood Rafiei
arXiv Computation and Language
Aug 27

Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models

The paper investigates how large language models (LLMs) perform multi‑hop reasoning and challenges the prevailing hop‑aligned circuit hypothesis, which posits that bridge entities are computed sequentially across layers. Through systematic analyses of real‑world multi‑hop queries, the authors discover a phenomenon called layer‑order inversion, where later‑hop answer entities become decodable earlier than bridge entities, and this effect grows with the number of hops. They propose a probabilistic recall‑and‑extract framework that models multi‑hop reasoning as broad probabilistic recall in shallow MLP layers followed by selective extraction in deeper attention layers, and validate this framework with probing analyses that reinterpret prior evidence, explain chain‑of‑thought gains, and diagnose multi‑hop failures.

By Xukai Liu, Ye Liu, Jipeng Zhang, Yanghai Zhang, Kai Zhang, Qi Liu