The paper proposes a method for automated research‑idea generation that preserves the typed structure of scientific papers by modeling each paper as a small category with typed research entities as objects and asserted relations as morphisms. It introduces a three‑layer algorithm—categorical signature clustering, a functor‑preservation gate, and a six‑axis LLM plausibility judge—to identify cross‑domain analogies that maintain relation chains. Experiments on tens of thousands of papers show the categorical gate filters candidates at a 17:1 ratio while keeping a falsifier rate above 83%, and it logs rejected candidates with detailed rationale.
By Yuchen Wang, Zhongzhi Luan
arXiv:2608.03729v4 Announce Type: replace-cross
Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large lang...
By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
The paper introduces GPTKB 2.0, a method for building disambiguated knowledge bases directly from large language models. It addresses the lack of native entity representation in LLMs by performing on‑the‑fly disambiguation of entities, relations, and classes, achieving a million‑scale KB with over 1 million disambiguated entities and 38.4 million triples. The authors analyze trade‑offs among accuracy, scale, and cost, and release the system at https://gptkb.org/.
By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
arXiv:2609.39572v1 Announce Type: new
Abstract: We analyze how lexical ambiguity and underspecification affect language model training. We create artificial homonyms and artificial hypernyms as pseud...
By Michaela Regneri, Nina Scheller, S\"oren Laue
arXiv:2608. 03729v1 Announce Type: cross Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source.
By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
arXiv:2410.17355v4 Announce Type: replace
Abstract: Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity t...
By Advait Deshmukh, Ashwin Umadi, Dananjay Srinivas, Maria Leonor Pacheco