arXiv Computer Vision

ReART: Reference-Guided Retrieval and Refinement for Emotion-Aware Art Generation

arXiv AI
2d ago

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

ATP‑Bench proposes a new benchmark for evaluating agentic tool planning in multimodal large language models (MLLMs) that generate interleaved text-and-image responses. The benchmark contains 7,702 QA pairs, including 1,592 visual‑question‑answer pairs, across eight categories and 25 visual‑critical intents, all verified by humans. A Multi‑Agent MLLM‑as‑a‑Judge (MAM) system is introduced to assess tool‑call precision, missed opportunities, and overall response quality without relying on ground‑truth references.

By Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang
Hugging Face Trending Papers
Aug 3

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion.

arXiv AI
Aug 7

CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

arXiv:2603. 08652v2 Announce Type: replace Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning.

By Haodong Li, Chunmei Qing, Huanyu Zhang, Dongzhi Jiang, Yihang Zou, Hongbo Peng, Dingming Li, Yuhong Dai, ZePeng Lin, Juanxi Tian, Yi Zhou, Siqi Dai, Jingwei Wu, Pheng-Ann Heng
Hugging Face Trending Papers
Jul 23

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases.

arXiv Computer Vision
3d ago

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.

By Long-Bao Nguyen, Quang-Khai Le, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
arXiv Computer Vision
6d ago

ID-VTG: Image-Disambiguated Video Temporal Grounding

arXiv:2608. 20127v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual attributes that are difficult to describe accurately in words alone.

By Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu
arXiv Computer Vision
2d ago

VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning

VinciCoder is a unified framework for multimodal code generation that addresses the limitations of single-task models by training on a large-scale curated corpus of 1.3 M direct generation pairs and 300 k visual‑refinement tasks. It introduces a coarse‑to‑fine Visual Reinforcement Learning (ViRL) approach that uses visual similarity across multi‑scale patches to provide an implementation‑agnostic reward, improving alignment between rendered outputs and input visuals. Experiments on diverse benchmarks show VinciCoder outperforms existing methods, and ablation studies confirm the effectiveness of ViRL.

By Xuanle Zhao, Deyang Jiang, Zhixiong Zeng, Lei Chen, Haoyue Yang, Haibo Qiu, Jing Huang, Yufeng Zhong, Liming Zheng, Yilin Cao, Lin Ma