arXiv Computation and Language

REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

The REAP system tackles the AKBC Shared Task 2026, aiming to build knowledge bases from language models without fine‑tuning and within a 32‑B parameter budget. It uses structured chain‑of‑thought reasoning, relation‑specific queries, and a reasoning‑based empty‑set gate to elicit knowledge, then directly extracts it into valid JSON arrays. Evaluated on the test set with the Mistral‑Small‑24B‑Instruct‑2501 model, REAP achieves a macro‑F1 of 0.62, notably high scores on countryLandBordersCountry (0.95), companyTradesAtStockExchange (0.73), and hasArea (0.77).

arXiv AI
Jul 28

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

arXiv:2607. 22639v1 Announce Type: new Abstract: Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search.

By Sai Shruthi Sistla, Ashutosh Hathidara, Christopher Toukmaji, Mayank Shrivastava, Karthikeyan Asokkumar
arXiv AI
Jun 2

KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models

arXiv:2604. 17621v2 Announce Type: replace Abstract: Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg.

By Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang, Johan Bos
arXiv Computation and Language
Sep 25

CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models

CONSISTRE is a consistency‑aware framework for document‑level relation extraction that tackles contradictions in large language model predictions. It offers two tracks: an inference‑time track that refines black‑box LLM outputs through constraint‑aware prompting, verification, and self‑reflection, and a training‑time track that distills consistency knowledge into smaller open‑source models via supervised fine‑tuning and reinforcement learning. Experiments on DocRED show both tracks outperform baselines, with the inference‑time track matching competitive F1 scores and the training‑time track narrowing the performance gap to proprietary LLMs while reducing inference cost.

By Mingxuan Sun
arXiv Computation and Language
Aug 25

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided proced...

By Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu
arXiv AI
2d ago

Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying

Build2SPARQL is a large-scale benchmark dataset for translating natural-language questions into SPARQL queries over building knowledge graphs. The dataset is generated by a KG‑grounded pipeline that produces 6,136 executable SPARQL queries and 30,680 corresponding natural-language questions across six query-pattern families and five vocabulary registers, covering 201 building KGs. Human validation shows high semantic fidelity, naturalness, and operational plausibility, and retrieval‑augmented evaluation demonstrates significant accuracy gains for open‑weight language models.

By Wooyoung Jung
arXiv Computation and Language
Sep 18

Schema-Anchored Latent Reasoning for Semantic Parsing-Based Knowledge Base Question Answering

The paper introduces SALR, a schema‑anchored latent reasoning approach for generating logical forms in knowledge‑base question answering. SALR delays explicit schema commitments by generating continuous thoughts in hidden states and aligns these thoughts with a codebook of KB schema elements, guided by an alignment objective derived from gold logical forms. Experiments on GrailQA and WebQSP demonstrate that SALR consistently outperforms strong baselines, notably improving compositional question performance by 2.86 F1 points over TIARA.

By Guangze Gao, Zixuan Li, Sikui Zhang, Chunfeng Yuan, Wenjuan Li, Bing Li, Xiaolong Jin, Weiming Hu