arXiv AI By Miguel Arana-Catania, Neil Jefferies

Characterising AI Models for Cataloguing

Read the original on arXiv AI →

arXiv:2607. 11353v1 Announce Type: cross Abstract: The creation of digital collections involves not only the digitisation of content, but also the creation of catalogue records for it.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 24

Ontology-supported AI Model and Dataset Management

The paper introduces an ontology-supported platform designed to facilitate the exchange, usage, and analysis of AI models and datasets. It addresses the need for effective management of AI assets in industrial settings by providing a structured framework that reduces semantic gaps. A real‑time critical systems use case demonstrates the platform’s practical utility.

By Jan Novacek, Ali Ahari, Tobias M\"uller, Sebastian Reiter, Alexander Viehl, Oliver Bringmann
arXiv Machine Learning
Sep 2

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

Modelpedia is an automated, LLM-assisted framework that extracts and organizes findings about AI models from published papers into a searchable public catalog. It links each finding to the relevant model, dataset, method, and concept, and has already extracted over a thousand findings from ICLR 2024 and 2025 papers. The authors invite the community to explore, contribute to, and build on this open catalog, positioning model findings as a shared foundation for the meta‑science of AI.

By Franciszek Bernat (Centre for Credible AI, Warsaw University of Technology), Dawid P{\l}udowski (Centre for Credible AI, Warsaw University of Technology), Micha{\l} Jan W{\l}odarczyk (Centre for Credible AI, Warsaw University of Technology), Luca Longo (University College Cork), Jianlong Zhou (University of Technology Sydney), Andreas Holzinger (Human-Centered AI Lab), Riccardo Guidotti (University of Pisa, ISTI-CNR), Wojciech Samek (Technical University of Berlin, Berlin Institute for the Foundations of Learning and Data), Przemys{\l}aw Biecek (Centre for Credible AI, University of Warsaw)
arXiv AI
Aug 24

TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding

TRACE is a new framework that uses agentic Large Language Models to automatically enrich e-commerce product catalogs with missing or buried attributes. It employs a ScoutAgent to gather multimodal evidence from merchant catalogs, syndicated feeds, and web search, and a JudgeAgent to verify and publish the proposed attribute values. In offline evaluation, TRACE achieved 98.2% accuracy with 74.7% coverage, and in production it increased enrichment coverage by 90.4% and boosted checkout conversion by 0.48%.

By Rohan Kumar, Steven Xu, Kyle MacDonald, Matthew Long, Bernice Chow, Mac VanRenterghem, Sudeep Das
arXiv AI
Sep 7

Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

The paper reports on building a research-software catalog using a coding agent, starting from a three‑day hackathon prototype and moving to public deployment. It details the engineering work needed—adversarial review, data‑quality checks, browser validation, and publication safeguards—to ensure reliable operation, noting that silent failures were more problematic than crashes. The authors then examine applying these lessons to a larger, human‑curated portal (MateriApps) that combines curated metadata, external documentation, vector search, and local language‑model generation, finding that explicit validation, monitoring, and repeated review remain essential for AI‑assisted software portals.

By Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada
arXiv Machine Learning
Jun 25

Autodata: An agentic data scientist to create high quality synthetic data

arXiv:2606. 25996v1 Announce Type: cross Abstract: We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data.

By Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston