Agents-K1: Towards Agent-native Knowledge Orchestration
arXiv:2606. 13669v1 Announce Type: new Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration.
arXiv:2608. 15193v1 Announce Type: cross Abstract: As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-term memory have improved information retrieval and task continuity.
arXiv:2606. 13669v1 Announce Type: new Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration.
Accumulated scientific knowledge advances inquiry when prior findings help researchers choose new questions, design investigations, and interpret results. Realizing this value at scale requires access...
arXiv:2607. 02609v1 Announce Type: cross Abstract: For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving organizational data.
The paper introduces the Large Knowledge Model (LKM), a scientific knowledge infrastructure that converts research literature into shared, computationally accessible reasoning graphs. LKM aligns questions, claims, and reasoning chains across papers, creating a Scientific Reasoning Landscape with Question, Workflow, and Evidence views. The system enhances scientific search, evidence‑grounded QA, and research planning, achieving notable accuracy gains on ChemBench, PubMedQA, and SciBench.
arXiv:2609.27297v2 Announce Type: replace Abstract: Agentic science envisions many autonomous agents investigating concurrently while building on a shared, evolving body of scientific knowledge. This...
The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.
The paper introduces Knowledge Cards, a new structured artefact designed to capture validated knowledge about specific concepts that AI systems use to make decisions. Unlike existing model, data, and system cards, Knowledge Cards focus on the layer between inputs and outputs, documenting entities, relationships, reasoning patterns, conditions for validity, and provenance, all grounded in a formal domain ontology and signed off by a domain expert. Prototype cards have been created in the energy and pharmaceutical domains, and the schema is released as a public draft for community engagement.
OptimusKG is a multimodal biomedical labeled property graph that integrates structured and semi‑structured resources to preserve detailed, type‑specific metadata across molecular, anatomical, clinical, and environmental domains. The graph contains nearly 191,000 nodes, over 21.8 million edges, and more than 67 million property instances derived from 18 ontologies, with a top‑level schema that enforces node and edge constraints while retaining granular provenance. Validation using the PaperQA3 agent found that 70.0% of sampled edges are supported by literature evidence, and the graph offers a standardized resource for machine learning, knowledge‑grounded retrieval, and hypothesis generation in biomedical research.
arXiv:2608. 14228v1 Announce Type: new Abstract: Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links.
arXiv:2607. 24512v1 Announce Type: new Abstract: Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles.
arXiv:2607. 21327v1 Announce Type: cross Abstract: Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI) analytics, yet suffer from temporal lag, semantic shallowness, and an inability to capture the non-linear dynamics of contemporary knowledge ecosystems.
arXiv:2608.29612v1 Announce Type: new Abstract: Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this...