CARAT: Do Materials LLMs Reason or Recite?
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper investigates how different components of a graph retrieval‑augmented generation pipeline affect large language model performance on knowledge‑graph question answering. It examines four variables—whether the answer path is included, the syntax of triples, the order of triples, and the subgraph size—across six LLMs and two benchmarks. The study finds that including the answer path is crucial, while the grounding instruction dramatically reduces accuracy when no facts are provided, and that syntax, order, and subgraph size have negligible measurable impact at multi‑hop depth.
arXiv:2609.22939v1 Announce Type: cross Abstract: Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for...
arXiv:2609.18154v1 Announce Type: cross Abstract: We describe our system for LitTraceQA (GroundLM @ EMNLP 2026): given a research question, retrieve the relevant papers from a pool of 27,487, cite th...
arXiv:2607. 28576v1 Announce Type: cross Abstract: Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought.
arXiv:2608. 07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment.
arXiv:2609.13267v1 Announce Type: new Abstract: Scientific charts encode quantities in axes, legends, and geometric marks, yet large vision-language models still treat them as natural photographs. Vi...