arXiv AI By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski

GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models

Read the original on arXiv AI →

arXiv:2608. 03729v1 Announce Type: cross Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Direct Construction of Disambiguated Knowledge Bases from Large Language Models

The paper introduces GPTKB 2.0, a method for building disambiguated knowledge bases directly from large language models. It addresses the lack of native entity representation in LLMs by performing on‑the‑fly disambiguation of entities, relations, and classes, achieving a million‑scale KB with over 1 million disambiguated entities and 38.4 million triples. The authors analyze trade‑offs among accuracy, scale, and cost, and release the system at https://gptkb.org/.

By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
arXiv Computation and Language
Sep 23

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

BELXTR is a new biomedical entity linking model that uses a multi‑vector (late interaction) architecture to preserve token‑level matching information, unlike traditional embedding‑based approaches that compress mentions into a single vector. By extending the XTR model with a task‑specific training objective and active query expansion, BELXTR achieves state‑of‑the‑art performance on half of ten evaluated corpora, with an average 5‑percentage‑point gain in recall@1. The model shows especially strong results on cross‑species gene disambiguation, outperforming an LLM‑powered retrieve‑and‑rerank pipeline and approaching a specialized rule‑based system.

By Samuele Garda, Ulf Leser
arXiv Computation and Language
Aug 31

Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection

The paper investigates modular entity disambiguation by separating candidate retrieval from entity selection. It compares sparse retrieval (BM25), Web KB search, and a dense retriever, all paired with large language model selectors. Results show that a training‑free BM25 retriever combined with an LLM selector achieves state‑of‑the‑art performance on the ZELDA benchmark, and the modular approach enables abstention when retrieval fails.

By Fina Polat, Daniel Daza, Pengyu Zhang, Klim Zaporojets, Paul Groth
arXiv AI
Sep 3

BioELX: Context-Aware Cross-lingual Biomedical Entity Linking without Task-Specific Supervision

BioELX is a retrieve‑rerank framework for cross‑lingual biomedical entity linking that tackles two key problems: the English‑biased UMLS alias training data and the degradation caused by naïvely adding context. It fine‑tunes SapBERT_multi with Wikidata‑derived cross‑lingual alias supervision to create shared concept neighborhoods, and then reranks candidates using pretrained LLMs with mention‑anchored prompting to focus on the target mention. Experiments demonstrate state‑of‑the‑art performance on four benchmarks, improving Recall@1 by 4.8–18.2 percentage points without task‑specific annotations.

By Yi Wang, Corina Dima, Liangyu Zhong, Steffen Staab