OpenAI Blog

Discovering types for entity disambiguation

Read the original on OpenAI Blog →

We’ve built a system for automatically figuring out which object is meant by a word by having a neural network decide if the word belongs to each of about 100 automatically-discovered “types” (non-exclusive categories).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

arXiv AI
Aug 24

Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

The paper proposes a method for automated research‑idea generation that preserves the typed structure of scientific papers by modeling each paper as a small category with typed research entities as objects and asserted relations as morphisms. It introduces a three‑layer algorithm—categorical signature clustering, a functor‑preservation gate, and a six‑axis LLM plausibility judge—to identify cross‑domain analogies that maintain relation chains. Experiments on tens of thousands of papers show the categorical gate filters candidates at a 17:1 ratio while keeping a falsifier rate above 83%, and it logs rejected candidates with detailed rationale.

By Yuchen Wang, Zhongzhi Luan
arXiv AI
Sep 3

Direct Construction of Disambiguated Knowledge Bases from Large Language Models

The paper introduces GPTKB 2.0, a method for building disambiguated knowledge bases directly from large language models. It addresses the lack of native entity representation in LLMs by performing on‑the‑fly disambiguation of entities, relations, and classes, achieving a million‑scale KB with over 1 million disambiguated entities and 38.4 million triples. The authors analyze trade‑offs among accuracy, scale, and cost, and release the system at https://gptkb.org/.

By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski