EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
SkillIR is a skill-guided framework for agentic image restoration that represents restoration experience as degradation-centered action evidence rather than full tool-use trajectories. It consolidates context-dependent action outcomes into scene-aware restoration skills, guiding one bounded action at a time within a verified residual-state loop. Experiments on synthetic and real-world multi-degradation datasets show that SkillIR improves restoration quality and enables more reliable and effective tool use.
arXiv:2608.29037v1 Announce Type: cross Abstract: Real-world degradations such as blur, shadow, distortion, and moire patterns severely impair the document question-answering capabilities of Multimod...
WeAgent-MMGenEdit is a comprehensive framework for multimodal agentic image generation and editing that addresses the unreliability of current models when prompts require external world knowledge. It introduces a multimodal harness with persistent evidence management, a scalable data construction pipeline producing 23K supervised trajectories and 14.7K RL tasks, and a bilingual benchmark (WeBench-MMGenEdit) for knowledge-intensive generation and multi-image editing. Post‑training methods based on SFT and RL further refine the agent policy and image backend, enabling a 30B‑parameter policy to outperform similarly sized models and approach the performance of a 1T‑parameter agent.
The paper introduces ImIR, a method that tunes a large pretrained image‑editing model for all‑in‑one image restoration by replacing text prompts with continuous image‑derived instructions. The approach uses a lightweight token mapper to shift the degraded image’s vision‑language embedding toward that of a clean image, enabling a single adapter to handle six restoration tasks in about three hours on one GPU. ImIR outperforms text conditioning in matched comparisons and supports task‑agnostic restoration without requiring a degradation label.
Degradations vary widely across images, so a practical restoration system has to handle many degradation types with one model. A recent and effective recipe adapts a large pretrained image-editing mod...
arXiv:2606. 01803v1 Announce Type: new Abstract: The explosive growth of Text-to-Image (T2I) models, from large-scale versions to lightweight, real-time ones, now faces diminishing marginal returns from single-model scaling.