MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs
arXiv:2507. 20804v3 Announce Type: replace Abstract: Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge.
arXiv:2608. 05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images.
arXiv:2507. 20804v3 Announce Type: replace Abstract: Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge.
arXiv:2607. 19128v1 Announce Type: new Abstract: Vision-language models (VLMs) provide a unified representation space for textual and visual information, yet their potential as general-purpose backbones for graph-structured data remains largely unexplored.
arXiv:2505.03654v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models have shown strong performance across multimodal tasks, and recent personalized MLLMs can recognize user-spec...
arXiv:2607. 15592v1 Announce Type: new Abstract: Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues.
PEARL is a new framework for inductive knowledge graph completion that treats relational paths as context-conditioned reasoning signals. It builds a query‑specific contextual subgraph from the query entities’ neighborhoods and uses a large language model‑guided retriever to select semantically relevant paths. By constructing a bipartite interaction graph over paths, contextual entities, and a global subgraph representation, and applying a dual‑view contrastive objective, PEARL adapts path embeddings to local and global structural evidence, achieving the best average Hits@10 on WN18RR, FB15k‑237, and NELL‑995.
arXiv:2511. 17731v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs).
VISPATH is a visual‑intent‑guided path reasoning framework designed for multimodal knowledge graph question answering (MM‑KGQA). It first identifies a reliable starting entity by fusing multimodal grounding with graph‑structural cues, then iteratively discovers and refines reasoning paths using hop‑specific multimodal intent and a reasoning‑chain pruning step. The framework is evaluated on the newly introduced VISPATH‑Bench, which tests two‑to‑four‑hop reasoning, and demonstrates consistent improvements over strong baselines, even surpassing GPT‑5.4 when using GPT‑4o as the backbone.
arXiv:2604. 04969v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) mitigates hallucinations in Multimodal Large Language Models (MLLMs), yet existing systems struggle with complex cross-modal reasoning.
arXiv:2609.05518v1 Announce Type: cross Abstract: Despite the strong capabilities of multimodal large language models (MLLMs), their parametric knowledge remains incomplete and difficult to update, m...
arXiv:2506. 02568v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated substantial efficacy in advancing graph-structured data analysis.
arXiv:2608. 15056v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or aggressively expanded evidence graphs, which can introduce noisy evidence, weaken multi-hop reasoning, and increase unsupported generation.
arXiv:2609.35942v1 Announce Type: new Abstract: Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs i...