FlyAOC is a benchmark that tests AI agents on end‑to‑end ontology curation of Drosophila scientific literature. Given a gene symbol, a brief description, a large paper corpus, and ontology resources, agents must search for evidence and produce structured annotations such as function terms, expression patterns, and historical synonyms. The benchmark contains 7,397 expert‑curated annotations across 100 genes and evaluates different agent harnesses, revealing system‑level failure modes that single‑task evaluations miss.
By Xingjian Zhang, Sophia Moylan, Ziyang Xiong, Qiaozhu Mei, Yichen Luo, Jiaqi W. Ma
arXiv:2609.00228v1 Announce Type: new
Abstract: Scientific domain entity linking (EL) differs from general domain EL because mentions and entity names often lack lexical overlap. Another challenge is...
By Md Rasel Khondokar, Qiao Qiao, Farjana Sultana Samia, Nhat Le, Yuepei Li, Qi Li
Padàrtha is the first ontology‑grounded fine‑grained Named Entity Recognition benchmark for Classical Sanskrit, built on the Mahêbhárata epic. Its tag set, derived from the Nyáya‑Vai’séka ontological system, contains 18 fine‑grained categories under 10 nodes and maps to five standard coarse tags, ensuring compatibility with existing benchmarks. The dataset includes over 12.6K expert‑annotated entries and 108,335 entity mentions across 73,632 verses, plus a 5,000‑verse test set designed to challenge rare mentions, and the study compares generative NER models to traditional architectures, noting performance drops at finer granularity and difficulties with unseen entities.
By Sujoy Sarkar, Pretam Ray, Paramhans Shah, Manoj Balaji Jagadeeshan, Akash Gairola, Arjuna S R, Pawan Goyal
arXiv:2609.10055v1 Announce Type: cross
Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical dat...
By Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang, Bairong Shen
Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical va...
arXiv:2608.31118v1 Announce Type: new
Abstract: The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled eval...
By Hamed Babaei Giglou, S\"oren Auer, Jennifer D'Souza