arXiv AI

Characterising AI Models for Cataloguing

arXiv:2607. 11353v1 Announce Type: cross Abstract: The creation of digital collections involves not only the digitisation of content, but also the creation of catalogue records for it.

arXiv AI
Aug 24

Ontology-supported AI Model and Dataset Management

The paper introduces an ontology-supported platform designed to facilitate the exchange, usage, and analysis of AI models and datasets. It addresses the need for effective management of AI assets in industrial settings by providing a structured framework that reduces semantic gaps. A real‑time critical systems use case demonstrates the platform’s practical utility.

By Jan Novacek, Ali Ahari, Tobias M\"uller, Sebastian Reiter, Alexander Viehl, Oliver Bringmann
arXiv Machine Learning
Sep 2

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

Modelpedia is an automated, LLM-assisted framework that extracts and organizes findings about AI models from published papers into a searchable public catalog. It links each finding to the relevant model, dataset, method, and concept, and has already extracted over a thousand findings from ICLR 2024 and 2025 papers. The authors invite the community to explore, contribute to, and build on this open catalog, positioning model findings as a shared foundation for the meta‑science of AI.

By Franciszek Bernat (Centre for Credible AI, Warsaw University of Technology), Dawid P{\l}udowski (Centre for Credible AI, Warsaw University of Technology), Micha{\l} Jan W{\l}odarczyk (Centre for Credible AI, Warsaw University of Technology), Luca Longo (University College Cork), Jianlong Zhou (University of Technology Sydney), Andreas Holzinger (Human-Centered AI Lab), Riccardo Guidotti (University of Pisa, ISTI-CNR), Wojciech Samek (Technical University of Berlin, Berlin Institute for the Foundations of Learning and Data), Przemys{\l}aw Biecek (Centre for Credible AI, University of Warsaw)
arXiv AI
Aug 24

TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding

TRACE is a new framework that uses agentic Large Language Models to automatically enrich e-commerce product catalogs with missing or buried attributes. It employs a ScoutAgent to gather multimodal evidence from merchant catalogs, syndicated feeds, and web search, and a JudgeAgent to verify and publish the proposed attribute values. In offline evaluation, TRACE achieved 98.2% accuracy with 74.7% coverage, and in production it increased enrichment coverage by 90.4% and boosted checkout conversion by 0.48%.

By Rohan Kumar, Steven Xu, Kyle MacDonald, Matthew Long, Bernice Chow, Mac VanRenterghem, Sudeep Das
arXiv AI
Sep 7

Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

The paper reports on building a research-software catalog using a coding agent, starting from a three‑day hackathon prototype and moving to public deployment. It details the engineering work needed—adversarial review, data‑quality checks, browser validation, and publication safeguards—to ensure reliable operation, noting that silent failures were more problematic than crashes. The authors then examine applying these lessons to a larger, human‑curated portal (MateriApps) that combines curated metadata, external documentation, vector search, and local language‑model generation, finding that explicit validation, monitoring, and repeated review remain essential for AI‑assisted software portals.

By Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada
arXiv Machine Learning
Jun 25

Autodata: An agentic data scientist to create high quality synthetic data

arXiv:2606. 25996v1 Announce Type: cross Abstract: We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data.

By Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston
arXiv AI
Sep 12

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

Benchmark Radar is a living database and search engine that aggregates AI benchmark papers, datasets, code, and score histories. It automatically discovers new benchmark resources from 37 sources, maintains a catalog of 1,283 records with 12,916 numeric observations, and provides tools such as a web dashboard, CLI, and downloadable evidence for researchers. The system also offers visualizations like a Pareto frontier and trend views to help users assess benchmark saturation and adoption.

By Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu
arXiv AI
Jun 29

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications

arXiv:2606. 28070v1 Announce Type: new Abstract: JD.

By Oxygen AIIC, Chan Long, Chao Liu, Chaofan Chen, Chaohui Dong, Chunyuan Guo, Danping Liu, Debin Liu, Deping Xiang, Fulai Xu, Guangyue Liu, Hao Li, Huichun Hu, Jian Yang, Jianan Wang, Jianbo Zhao, Jiaoyang Li, Jiaxing Wang, Jinglong Li, Jinjin Guo, Jun Fang, Jun Liu, Kai Zhou, Li Wang, Lili Gao, Liying Chen, Luning Yang, Mengdi Zhou, Pengzhang Liu, Qi Lv, Qianyun Wang, Qixia Jiang, Ruyue Li, Shimu Liang, Shuxing Wang, Sijie Zhang, Siqi Li, Tianhao Gao, Wang Ke, Weihu Huang, Wencan Lai, Wenjie Zhang, Xiaohui Zhang, Xiaojing Dong, Ya Liu, Yifeng Zhang, Yixiang Wang, Yongtai Zhang, Yongyi Liao, Zhaoru Chen, Zhen Chen, Zhiyong Ma, Zhiyuan Liu, Zhongwei Liu, Ziyan Xing