The paper presents an end‑to‑end pipeline for automatically constructing a tree‑structured knowledge graph (KG) from Vietnamese high school history textbooks and evaluating retrieval strategies that exploit the KG’s hierarchical structure. The KG construction uses a three‑phase hybrid relation extraction process, including intra‑batch deduplication, approximate cross‑batch search, and LLM extraction with a centroid filter and dual‑LLM validator, resulting in 750 nodes and 4,341 semantic edges across 41 ontology types. Retrieval evaluation compares three graph traversal strategies—Top‑Down, Horizontal, and Bottom‑Up—on a benchmark of 1,210 Vietnamese queries, finding that the Top‑Down strategy with structural information outperforms a vector baseline by 4.7 percentage points in NDCG@10.
By Ket Doan Nguyen, Minh N. H. Nguyen
arXiv:2608. 10644v1 Announce Type: new Abstract: Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions are destructive in a way extraction errors are not.
By Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.
By Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou
arXiv:2603. 06915v2 Announce Type: replace-cross Abstract: The extraction of structured information from raw text is a fundamental component of many NLP applications, including document retrieval, ranking, and relevance estimation.
By Moin Amin-Naseri, Hannah Kim, Estevam Hruschka
arXiv:2606. 19710v1 Announce Type: cross Abstract: Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargon-heavy legal documents.
By Elijah Feldman, Dipak Meher, Carlotta Domeniconi
arXiv:2607. 10212v1 Announce Type: new Abstract: Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance.
By Nipun Misra, Vikranth Udandarao, Aanchal Gupta, Yogender Kumar, Manuj Mukherjee, Raghava Mutharaju
Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-specific metrics or small-scale manual verification, offering limited insight into the structural and semantic fidelity of extracted graphs.
arXiv:2608. 14228v1 Announce Type: new Abstract: Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links.
By Yiming Zhang, Koji Tsuda
W-RAG is a source-aware retrieval framework designed for enterprise document generation from heterogeneous knowledge bases. It uses ontology-guided retrieval, local ranking within each knowledge base, and source-level weighting to balance evidence from diverse sources. A new dataset covering multiple document types and industry domains demonstrates that W-RAG improves document coverage and generation quality compared to standard RAG pipelines.
By Hridya Dhulipala, Rajesh Ombase, Michael Wang, Tien N. Nguyen
ConvergeWriter introduces a bottom‑up, data‑driven framework for long‑form document generation that first retrieves exhaustive knowledge from a source corpus and clusters it into distinct knowledge groups. These clusters then guide the creation of a hierarchical outline and the final text, ensuring the output is strictly grounded in the retrieved material and traceable to its sources. Experiments on 14B and 32B LLMs show that this approach matches or surpasses state‑of‑the‑art baselines, especially in scenarios requiring high factual fidelity and structural coherence.
By Binquan Ji, Jiaqi Wang, Ruiting Li, Xingchen Han, Yiyang Qi, Shichao Wang, Yifei Lu, Yuantao Han, Feiliang Ren
arXiv:2607. 16201v1 Announce Type: new Abstract: Ontology engineering remains a critical bottleneck in knowledge-intensive AI systems.
By Sergei Sergienko
CORTEX is a novel framework that transforms web‑scale corpus construction from flat document filtering into structured knowledge organization using an Ontological Corpus Graph (OCG). The OCG comprises a quality‑refined content layer, a lightweight ontology layer that evolves via LLMs, and a cross‑domain alignment layer that supports arbitrary taxonomic resolution. Experiments demonstrate CORTEX’s effectiveness, and the authors release a 24.14 B‑token refined corpus, its OCG, and a cross‑domain benchmark called CortexBench for evaluating large language models.
By Chengtao Gan, Xiaoke Guo, Yushan Zhu, Zhaoyan Gong, Zhiqiang Liu, Songze Li, Huajun Chen, Wen Zhang