The paper investigates how different components of a graph retrieval‑augmented generation pipeline affect large language model performance on knowledge‑graph question answering. It examines four variables—whether the answer path is included, the syntax of triples, the order of triples, and the subgraph size—across six LLMs and two benchmarks. The study finds that including the answer path is crucial, while the grounding instruction dramatically reduces accuracy when no facts are provided, and that syntax, order, and subgraph size have negligible measurable impact at multi‑hop depth.
By Arquimedes Canedo
arXiv:2609.22939v1 Announce Type: cross
Abstract: Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for...
By Wenji Fu
arXiv:2609.18154v1 Announce Type: cross
Abstract: We describe our system for LitTraceQA (GroundLM @ EMNLP 2026): given a research question, retrieve the relevant papers from a pool of 27,487, cite th...
By Aaditya Chauhan
arXiv:2607. 28576v1 Announce Type: cross Abstract: Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought.
By Iliya Mirzaei
arXiv:2608. 07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment.
By Zhengze Huang, Luyang Yu, Di Hong, Xinzhe Huang, Wanyu Lin, Zhixuan Chu, Zhan Qin, Tianhang Zheng
arXiv:2609.13267v1 Announce Type: new
Abstract: Scientific charts encode quantities in axes, legends, and geometric marks, yet large vision-language models still treat them as natural photographs. Vi...
By Alberlucia Rafael Soarez, Camila Ferreira, Daniel Kim, Mariana Costa, Alejandro Torres
arXiv:2609.39142v1 Announce Type: new
Abstract: Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information th...
By Yunbei Zhang, Janet Wang, Jihun Hamm, Chandan K Reddy
arXiv:2608. 08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said.
By Fengrong Wan, Chengcan Wu, Ningtao Lyu
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
arXiv:2609.01556v1 Announce Type: cross
Abstract: We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not...
By Nabira Rashid, Manolis Kellis
arXiv:2607. 14552v1 Announce Type: cross Abstract: A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors.
By Jungseob Lee, Seungyoon Lee, Suhyune Son, Dongyub Jude Lee, Sungbin Han, Sugyeong Eo, Heuiseok Lim
arXiv:2608. 18242v1 Announce Type: new Abstract: We introduce ClosureBench, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth.
By Stefano Goria (AIM Research Lab)