arXiv Computer Vision

EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

arXiv Computer Vision
Sep 21

SkillIR: Evolving Scene-Aware Skills for Agentic Image Restoration

SkillIR is a skill-guided framework for agentic image restoration that represents restoration experience as degradation-centered action evidence rather than full tool-use trajectories. It consolidates context-dependent action outcomes into scene-aware restoration skills, guiding one bounded action at a time within a verified residual-state loop. Experiments on synthetic and real-world multi-degradation datasets show that SkillIR improves restoration quality and enables more reliable and effective tool use.

By Jie Shao, Shengkai Hu, Xu Zhang, Beihang Song, Yongcheng Jing, Xu Wu, Jun Wan
arXiv Computer Vision
Sep 7

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

WeAgent-MMGenEdit is a comprehensive framework for multimodal agentic image generation and editing that addresses the unreliability of current models when prompts require external world knowledge. It introduces a multimodal harness with persistent evidence management, a scalable data construction pipeline producing 23K supervised trajectories and 14.7K RL tasks, and a bilingual benchmark (WeBench-MMGenEdit) for knowledge-intensive generation and multi-image editing. Post‑training methods based on SFT and RL further refine the agent policy and image backend, enabling a 30B‑parameter policy to outperform similarly sized models and approach the performance of a 1T‑parameter agent.

By Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu, Han Li, Zhen Cao, Wenchao Chen, Chengduo Zhao, Fandong Meng
arXiv Computer Vision
Sep 23

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

The paper introduces ImIR, a method that tunes a large pretrained image‑editing model for all‑in‑one image restoration by replacing text prompts with continuous image‑derived instructions. The approach uses a lightweight token mapper to shift the degraded image’s vision‑language embedding toward that of a clean image, enabling a single adapter to handle six restoration tasks in about three hours on one GPU. ImIR outperforms text conditioning in matched comparisons and supports task‑agnostic restoration without requiring a degradation label.

By S\"uleyman Aslan, G\"orkay Aydemir, M{\i}sra Yavuz, Yunus Bilge Kurt, Nasrin Rahimi, Ahmet Rasim Emirda\u{g}{\i}, Burak Can Biner, M. Ak{\i}n Y{\i}lmaz
arXiv AI
3d ago

EviRover: Reinforcing Agentic Perception Beyond a Glance

EviRover is a perception agent that goes beyond a single glance by actively gathering information to resolve perceptual queries. The authors created two data generation pipelines, producing EviRover-SFT-5K and EviRover-RL-12K, and a human‑verified benchmark called EviLens with 688 instances across five perception categories. Trained with supervised fine‑tuning and agentic reinforcement learning, the 4B EviRover outperforms its backbone by an average of 30 points on EviLens and shows strong transfer to other benchmarks such as WebEyes and BrowseComp‑VL.

By Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
Hugging Face Trending Papers
Jun 22

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous intentions, logical reasoning, and Out-of-Distribution (OOD) knowledge, existing image models often yield sub-optimal results due to a lack of deep reasoning capabilities and real-time external information.

arXiv Computation and Language
Aug 27

VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following

VISA (Visual Instruction Synthesis Agent) is an agentic framework that transforms multimodal instruction synthesis into a self‑evolving loop. Each cycle analyzes images to filter constraints, samples new constraint sets, generates candidate instructions, and verifies them using executable tools and large language model judges. Failed samples trigger diagnostic recovery, while accepted samples are evaluated against the target model to estimate difficulty, with all feedback written back to memory to adapt future rounds and provide reward signals for reinforcement learning.

By Min Zeng, Guanxin Tan, Libin Cen, Yawei Wen, Rui Hu, Liuyang Bian, Xiaolong Chen, Xiaoxin Chen
Hugging Face Trending Papers
Aug 12

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes.