arXiv AI

DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA

arXiv:2510. 16302v2 Announce Type: replace Abstract: Multi-hop reasoning for question answering (QA) plays a critical role in retrieval-augmented generation (RAG) for modern large language models (LLMs).

arXiv Machine Learning
Aug 27

AlgoTrace: Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models

The paper introduces AlgoTrace, a framework that traces and steers algorithmic operations in large language models’ latent space during multi‑step reasoning. By clustering latent activations on tasks such as TSP, 3SAT, AIME, and Graph Navigation, the authors identify reusable primitive vectors that can be injected to elicit specific algorithmic behaviors, composed algebraically, and transferred across models and tasks. Fine‑tuning further improves the composition of these primitives, suggesting that LLM reasoning can be viewed as a walk through algorithmic primitives governed by compositional geometry.

By Samuel Lippl, Thomas McGee, Kimberly Lopez, Ziwen Pan, Pierce Zhang, Salma Ziadi, Oliver Eberle, Ida Momennejad
arXiv Machine Learning
Sep 21

VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

VISPATH is a visual‑intent‑guided path reasoning framework designed for multimodal knowledge graph question answering (MM‑KGQA). It first identifies a reliable starting entity by fusing multimodal grounding with graph‑structural cues, then iteratively discovers and refines reasoning paths using hop‑specific multimodal intent and a reasoning‑chain pruning step. The framework is evaluated on the newly introduced VISPATH‑Bench, which tests two‑to‑four‑hop reasoning, and demonstrates consistent improvements over strong baselines, even surpassing GPT‑5.4 when using GPT‑4o as the backbone.

By Jinke Wu, Zhengpin Li, Mengzhe Jia, Yang Li, Wentao Zhang
arXiv AI
Aug 26

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

arXiv:2603.16654v3 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, espe...

By Xiaojie Gu, Sherry T. Tong, Aosong Feng, Sophia Simeng Han, Jinghui Lu, Yingjian Chen, Yusuke Iwasawa, Yutaka Matsuo, Chanjun Park, Rex Ying, Irene Li