FlyAOC is a benchmark that tests AI agents on end‑to‑end ontology curation of Drosophila scientific literature. Given a gene symbol, a brief description, a large paper corpus, and ontology resources, agents must search for evidence and produce structured annotations such as function terms, expression patterns, and historical synonyms. The benchmark contains 7,397 expert‑curated annotations across 100 genes and evaluates different agent harnesses, revealing system‑level failure modes that single‑task evaluations miss.
By Xingjian Zhang, Sophia Moylan, Ziyang Xiong, Qiaozhu Mei, Yichen Luo, Jiaqi W. Ma
arXiv:2609.00228v1 Announce Type: new
Abstract: Scientific domain entity linking (EL) differs from general domain EL because mentions and entity names often lack lexical overlap. Another challenge is...
By Md Rasel Khondokar, Qiao Qiao, Farjana Sultana Samia, Nhat Le, Yuepei Li, Qi Li
Padàrtha is the first ontology‑grounded fine‑grained Named Entity Recognition benchmark for Classical Sanskrit, built on the Mahêbhárata epic. Its tag set, derived from the Nyáya‑Vai’séka ontological system, contains 18 fine‑grained categories under 10 nodes and maps to five standard coarse tags, ensuring compatibility with existing benchmarks. The dataset includes over 12.6K expert‑annotated entries and 108,335 entity mentions across 73,632 verses, plus a 5,000‑verse test set designed to challenge rare mentions, and the study compares generative NER models to traditional architectures, noting performance drops at finer granularity and difficulties with unseen entities.
By Sujoy Sarkar, Pretam Ray, Paramhans Shah, Manoj Balaji Jagadeeshan, Akash Gairola, Arjuna S R, Pawan Goyal
arXiv:2609.10055v1 Announce Type: cross
Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical dat...
By Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang, Bairong Shen
Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical va...
arXiv:2608.31118v1 Announce Type: new
Abstract: The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled eval...
By Hamed Babaei Giglou, S\"oren Auer, Jennifer D'Souza
arXiv:2608. 07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactically valid and semantically faithful.
By Tommaso Soru, Abdulsobur Oyewale
arXiv:2604. 08552v2 Announce Type: replace-cross Abstract: Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reuse.
By Josef Hardi, Martin J. O'Connor, Marcos Martinez-Romero, Jean G. Rosario, Stephen A. Fisher, Mark A. Musen
arXiv:2607. 01977v1 Announce Type: new Abstract: Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, and evaluation practices.
By Hamed Babaei Giglou, Jennifer D'Souza, Andrei Aioanei, Nandana Mihindukulasooriya, S\"oren Auer
arXiv:2608. 14228v1 Announce Type: new Abstract: Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links.
By Yiming Zhang, Koji Tsuda
arXiv:2610.01393v1 Announce Type: cross
Abstract: Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary...
By Nouha Hayouni, Sheeba Samuel, Alsayed Algergawy
arXiv:2608. 14587v1 Announce Type: new Abstract: Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance.
By Nicolas Turenne, Youcef Sklab, Eric Chenin, Jean-Daniel Zucker