Measuring Similarity between Artistic and AI Generated Images using Siamese Neural Networks
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces Preserve-and-Compose Training (PACT) for composed image retrieval, a task where a query image is modified by a textual instruction while preserving visual content from a reference image. PACT learns from image–text–text triplets, using target captions for supervision and visual evidence from the source image to maintain relevant details, without requiring target images or gallery updates. The authors also propose Chord scoring, which blends target similarity with source-relative directional agreement in a frozen image space, and demonstrate that this combined approach yields strong retrieval performance across multiple zero-shot CIR benchmarks and various backbones.
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
arXiv:2603. 01696v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) often omit or misrepresent critical visual content in generated image captions.
arXiv:2604. 11539v2 Announce Type: replace-cross Abstract: Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus.
The aim of this paper is twofold. First, it investigates whether newer generative models are getting better at pastiching contemporary artworks.
The paper explores whether CLIP embeddings can detect AI-generated images by using a frozen CLIP model to extract visual embeddings and training lightweight classifiers on top. On the CIFAKE benchmark, the approach achieves 95% accuracy without language reasoning, and 85% accuracy after few-shot adaptation with 20% of the data. Certain image types, such as wide-angle photographs and oil paintings, remain challenging, highlighting unexplored difficulties in AI-generated image classification.