arXiv:2607. 02609v1 Announce Type: cross Abstract: For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving organizational data.
By Mariano Garralda-Barrio
arXiv:2608. 15193v1 Announce Type: cross Abstract: As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-term memory have improved information retrieval and task continuity.
By Yuyang Zheng, Nan Li, Wenxia Deng, Lige Yan, Xiang Li, Si Chen
arXiv:2608. 08056v1 Announce Type: new Abstract: Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals.
By Ioannis N. Tzortzis, Georgia Kapetadimitri, Agapi Davradou, Nefeli Kousta, Nikolaos Bakalos, Ioannis Rallis, Dimitrios Kalogeras, Nikolaos Doulamis, Anastasios Doulamis
The paper introduces an architectural mediation approach that uses the Model Context Protocol (MCP) to bridge large language model (LLM) agents with data spaces. By implementing the Eunomia Agent, the mediation layer translates data space capabilities into structured, schema-driven tools that LLM agents can discover and invoke while respecting governance constraints. A prototype demonstrates end‑to‑end interaction across catalog discovery, metadata retrieval, and data service invocation without altering existing data space components, showing that protocol‑based mediation enables interoperable, standards‑aligned integration of AI agents into governed data‑sharing ecosystems.
By Jaime Alonso Ruiz, Carlos Aparicio, Gabriel Huecas, Joaqu\'in Salvach\'ua, Andres Munoz-Arcentales
The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.
By Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou
arXiv:2607. 19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows.
By Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun