arXiv AI

OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models

arXiv:2607. 01977v1 Announce Type: new Abstract: Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, and evaluation practices.

arXiv Computation and Language
Sep 18

CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

CORTEX is a novel framework that transforms web‑scale corpus construction from flat document filtering into structured knowledge organization using an Ontological Corpus Graph (OCG). The OCG comprises a quality‑refined content layer, a lightweight ontology layer that evolves via LLMs, and a cross‑domain alignment layer that supports arbitrary taxonomic resolution. Experiments demonstrate CORTEX’s effectiveness, and the authors release a 24.14 B‑token refined corpus, its OCG, and a cross‑domain benchmark called CortexBench for evaluating large language models.

By Chengtao Gan, Xiaoke Guo, Yushan Zhu, Zhaoyan Gong, Zhiqiang Liu, Songze Li, Huajun Chen, Wen Zhang
arXiv AI
Aug 28

pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning

The pro-team at LLMs4OL 2026 presented a system for ontology learning that tackles both the End-to-End Flagship Task (Task A) and the Ontology Extension Reuse Task (Task B). Their approach uses an offline retrieval‑augmented few‑shot prompting pipeline with Qwen2.5‑14B‑Instruct and MiniLM‑L6‑v2 for retrieval, selecting top‑5 examples for Task A and top‑2 for Task B, and applies a left‑truncated context‑windowing strategy to keep task instructions in long prompts. For Task B, generated triples are filtered deterministically by a vocabulary constraint, keeping triples that involve at least one term from the closed vocabulary and removing duplicates of the initial ontology, achieving high scores in Semantic Graph Similarity, Term‑Typing F1, and Taxonomy Discovery F1, though no non‑taxonomic relations were extracted.

By Shivam Mishra, Dhannu Ram Meena, Muneendra Ojha, Krishna Pratap Singh, Kuldeep Singh
arXiv AI
Jul 28

An Ontology for Machine Learning Interatomic Potentials

arXiv:2607. 23219v1 Announce Type: new Abstract: Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density functional theory (DFT) or wave-function methods---at a fraction of the cost.

By Daniel Hern\'andez, Jong Hyun Jung, Yuji Ikeda, Yongliang Ou, Pranav Kumar, Tom Sch\"achtel, Wenchuan Liu, Xin Li, Xi Zhang, Xiang Xu, Lifang Zhu, Fritz K\"ormann, Steffen Staab, Blazej Grabowski
arXiv AI
Sep 10

Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?

The paper introduces Open Tabular Insight Extraction (OpenTI), a unified framework aimed at democratizing access to insights from large table corpora. It highlights how current research is fragmented across domains like table QA, text‑to‑SQL, and data analysis agents, and shows that existing systems and benchmarks fall short of covering the full end‑to‑end scope of OpenTI. The authors propose a consolidated terminology, conduct a systematic review, and outline a research agenda for developing comprehensive OpenTI systems, evaluation methods, and interaction paradigms.

By Daniel Gomm, Maarten de Rijke, Madelon Hulsebos
arXiv Computation and Language
Sep 10

OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

arXiv:2609.10055v1 Announce Type: cross Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical dat...

By Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang, Bairong Shen