Knowing When to Trust Images: Reliability-Aware Multi-modal Entity Alignment
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
RiVaT‑Fuse introduces a reliability‑calibrated variational tensor fusion framework for multimodal image‑metadata prediction, treating fusion as a sample‑wise latent‑state estimation rather than simple aggregation. It replaces scalar modality confidence with matrix‑valued trust geometry, decomposes interactions into additive, multiplicative, and relational components, and couples the latent state with conditional robustness and structured multi‑task prediction. On an image‑level benchmark, RiVaT‑Fuse outperforms direct representation‑level baselines and improves probability and label stability under perturbation.
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.
arXiv:2609.12397v1 Announce Type: new Abstract: Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years. Despite the rapid advancement...
arXiv:2507. 20804v3 Announce Type: replace Abstract: Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge.
arXiv:2606. 17950v1 Announce Type: cross Abstract: Visual information helps resolve ambiguity in coreference resolution, leading to notable performance gains.