arXiv AI

The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons

arXiv:2607. 00032v1 Announce Type: new Abstract: Many information systems are built around documents: self-contained units optimised for print production and linear reading.

arXiv AI
Aug 11

H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System

arXiv:2608. 08056v1 Announce Type: new Abstract: Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals.

By Ioannis N. Tzortzis, Georgia Kapetadimitri, Agapi Davradou, Nefeli Kousta, Nikolaos Bakalos, Ioannis Rallis, Dimitrios Kalogeras, Nikolaos Doulamis, Anastasios Doulamis
arXiv AI
6d ago

Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol

The paper introduces an architectural mediation approach that uses the Model Context Protocol (MCP) to bridge large language model (LLM) agents with data spaces. By implementing the Eunomia Agent, the mediation layer translates data space capabilities into structured, schema-driven tools that LLM agents can discover and invoke while respecting governance constraints. A prototype demonstrates end‑to‑end interaction across catalog discovery, metadata retrieval, and data service invocation without altering existing data space components, showing that protocol‑based mediation enables interoperable, standards‑aligned integration of AI agents into governed data‑sharing ecosystems.

By Jaime Alonso Ruiz, Carlos Aparicio, Gabriel Huecas, Joaqu\'in Salvach\'ua, Andres Munoz-Arcentales
arXiv AI
Sep 12

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.

By Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou
arXiv AI
Jul 23

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

arXiv:2607. 19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows.

By Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun
arXiv AI
Aug 28

Knowledge Cards: Structured Knowledge for AI Systems

The paper introduces Knowledge Cards, a new structured artefact designed to capture validated knowledge about specific concepts that AI systems use to make decisions. Unlike existing model, data, and system cards, Knowledge Cards focus on the layer between inputs and outputs, documenting entities, relationships, reasoning patterns, conditions for validity, and provenance, all grounded in a formal domain ontology and signed off by a domain expert. Prototype cards have been created in the energy and pharmaceutical domains, and the schema is released as a public draft for community engagement.

By Liliana Ferreira