Knowledge-Centric Information Systems
arXiv:2607. 02609v1 Announce Type: cross Abstract: For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving organizational data.
arXiv:2607. 00032v1 Announce Type: new Abstract: Many information systems are built around documents: self-contained units optimised for print production and linear reading.
arXiv:2607. 02609v1 Announce Type: cross Abstract: For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving organizational data.
arXiv:2608. 15193v1 Announce Type: cross Abstract: As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-term memory have improved information retrieval and task continuity.
arXiv:2608. 08056v1 Announce Type: new Abstract: Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals.
The paper introduces an architectural mediation approach that uses the Model Context Protocol (MCP) to bridge large language model (LLM) agents with data spaces. By implementing the Eunomia Agent, the mediation layer translates data space capabilities into structured, schema-driven tools that LLM agents can discover and invoke while respecting governance constraints. A prototype demonstrates end‑to‑end interaction across catalog discovery, metadata retrieval, and data service invocation without altering existing data space components, showing that protocol‑based mediation enables interoperable, standards‑aligned integration of AI agents into governed data‑sharing ecosystems.
The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.
arXiv:2607. 19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows.
arXiv:2608. 12133v1 Announce Type: new Abstract: Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images.
The paper introduces Knowledge Cards, a new structured artefact designed to capture validated knowledge about specific concepts that AI systems use to make decisions. Unlike existing model, data, and system cards, Knowledge Cards focus on the layer between inputs and outputs, documenting entities, relationships, reasoning patterns, conditions for validity, and provenance, all grounded in a formal domain ontology and signed off by a domain expert. Prototype cards have been created in the energy and pharmaceutical domains, and the schema is released as a public draft for community engagement.
arXiv:2607. 18029v1 Announce Type: cross Abstract: Researchers need to answer ad-hoc questions about the contents of domain-specific archives but often lack the expertise to write structured queries on the metadata.
arXiv:2604. 08552v2 Announce Type: replace-cross Abstract: Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reuse.
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific kno...
arXiv:2608.23263v1 Announce Type: new Abstract: The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to pro...