arXiv AI

FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

FlyAOC is a benchmark that tests AI agents on end‑to‑end ontology curation of Drosophila scientific literature. Given a gene symbol, a brief description, a large paper corpus, and ontology resources, agents must search for evidence and produce structured annotations such as function terms, expression patterns, and historical synonyms. The benchmark contains 7,397 expert‑curated annotations across 100 genes and evaluates different agent harnesses, revealing system‑level failure modes that single‑task evaluations miss.

arXiv AI
Jun 4

BRAINCELL-AID: An Agentic AI Created Brain Cell Type Resource for Community Annotation

arXiv:2510. 17064v4 Announce Type: replace Abstract: Single-cell RNA sequencing has transformed our ability to identify diverse cell types and their transcriptomic signatures.

By Rongbin Li, Wenbo Chen, Zhao Li, Rodrigo Munoz-Castaneda, Jinbo Li, Neha S. Maurya, Arnav Solanki, Huan He, Hanwen Xing, Meaghan Ramlakhan, Zachary Wise, Nelson Johansen, Zhuhao Wu, Hua Xu, Michael Hawrylycz, W. Jim Zheng
arXiv AI
Jun 12

Agents-K1: Towards Agent-native Knowledge Orchestration

arXiv:2606. 13669v1 Announce Type: new Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration.

By Zongsheng Cao, Bihao Zhan, Jinxin Shi, Jiong Wang, Fangchen Yu, Zhijie Zhong, Zijie Guo, Tianshuo Peng, Zhuo Liu, Yi Xie, Xiang Zhuang, Yue Fan, Runmin Ma, Shiyang Feng, Xiangchao Yan, Anran Liu, Peng Ye, Wenlong Zhang, Shufei Zhang, Chunfeng Song, Fenghua Ling, Jie Zhou, Liang He, Bo Zhang, Lei Bai
arXiv Machine Learning
Sep 2

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

Modelpedia is an automated, LLM-assisted framework that extracts and organizes findings about AI models from published papers into a searchable public catalog. It links each finding to the relevant model, dataset, method, and concept, and has already extracted over a thousand findings from ICLR 2024 and 2025 papers. The authors invite the community to explore, contribute to, and build on this open catalog, positioning model findings as a shared foundation for the meta‑science of AI.

By Franciszek Bernat (Centre for Credible AI, Warsaw University of Technology), Dawid P{\l}udowski (Centre for Credible AI, Warsaw University of Technology), Micha{\l} Jan W{\l}odarczyk (Centre for Credible AI, Warsaw University of Technology), Luca Longo (University College Cork), Jianlong Zhou (University of Technology Sydney), Andreas Holzinger (Human-Centered AI Lab), Riccardo Guidotti (University of Pisa, ISTI-CNR), Wojciech Samek (Technical University of Berlin, Berlin Institute for the Foundations of Learning and Data), Przemys{\l}aw Biecek (Centre for Credible AI, University of Warsaw)
arXiv AI
Sep 12

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.

By Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou
arXiv Computation and Language
Sep 1

Ontology-Guided Multi-Agent Extraction of Evaluation Objects from Academic Review Texts: Evidence from Chinese Library and Information Science

The paper introduces an ontology‑guided multi‑agent framework for extracting evaluation objects from academic review texts, addressing challenges such as abstractness, context‑dependency, and ambiguous type boundaries. The system combines candidate discovery, ontology‑constrained classification, and domain review, achieving high precision (90.33%) and recall (84.55%) and outperforming rule‑based and zero‑shot baselines. Ablation studies show that the multi‑agent workflow boosts recall and stability, while ontology‑based constraints improve fine‑grained classification and reduce category confusion.

By Haolin Chen, Hongyi Dong, Yu Zhu, Yijia Hong, Leiqing Niu, Jiyuan Ye