arXiv Computer Vision
Sep 21

SkillIR: Evolving Scene-Aware Skills for Agentic Image Restoration

SkillIR is a skill-guided framework for agentic image restoration that represents restoration experience as degradation-centered action evidence rather than full tool-use trajectories. It consolidates context-dependent action outcomes into scene-aware restoration skills, guiding one bounded action at a time within a verified residual-state loop. Experiments on synthetic and real-world multi-degradation datasets show that SkillIR improves restoration quality and enables more reliable and effective tool use.

By Jie Shao, Shengkai Hu, Xu Zhang, Beihang Song, Yongcheng Jing, Xu Wu, Jun Wan
arXiv Computer Vision
Sep 7

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

WeAgent-MMGenEdit is a comprehensive framework for multimodal agentic image generation and editing that addresses the unreliability of current models when prompts require external world knowledge. It introduces a multimodal harness with persistent evidence management, a scalable data construction pipeline producing 23K supervised trajectories and 14.7K RL tasks, and a bilingual benchmark (WeBench-MMGenEdit) for knowledge-intensive generation and multi-image editing. Post‑training methods based on SFT and RL further refine the agent policy and image backend, enabling a 30B‑parameter policy to outperform similarly sized models and approach the performance of a 1T‑parameter agent.

By Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu, Han Li, Zhen Cao, Wenchao Chen, Chengduo Zhao, Fandong Meng
arXiv Computer Vision
Sep 23

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

The paper introduces ImIR, a method that tunes a large pretrained image‑editing model for all‑in‑one image restoration by replacing text prompts with continuous image‑derived instructions. The approach uses a lightweight token mapper to shift the degraded image’s vision‑language embedding toward that of a clean image, enabling a single adapter to handle six restoration tasks in about three hours on one GPU. ImIR outperforms text conditioning in matched comparisons and supports task‑agnostic restoration without requiring a degradation label.

By S\"uleyman Aslan, G\"orkay Aydemir, M{\i}sra Yavuz, Yunus Bilge Kurt, Nasrin Rahimi, Ahmet Rasim Emirda\u{g}{\i}, Burak Can Biner, M. Ak{\i}n Y{\i}lmaz