arXiv AI By Hamed Babaei Giglou, Jennifer D'Souza, Andrei Aioanei, Nandana Mihindukulasooriya, S\"oren Auer

OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models

Read the original on arXiv AI →

arXiv:2607. 01977v1 Announce Type: new Abstract: Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, and evaluation practices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 18

CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

CORTEX is a novel framework that transforms web‑scale corpus construction from flat document filtering into structured knowledge organization using an Ontological Corpus Graph (OCG). The OCG comprises a quality‑refined content layer, a lightweight ontology layer that evolves via LLMs, and a cross‑domain alignment layer that supports arbitrary taxonomic resolution. Experiments demonstrate CORTEX’s effectiveness, and the authors release a 24.14 B‑token refined corpus, its OCG, and a cross‑domain benchmark called CortexBench for evaluating large language models.

By Chengtao Gan, Xiaoke Guo, Yushan Zhu, Zhaoyan Gong, Zhiqiang Liu, Songze Li, Huajun Chen, Wen Zhang