EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images.
arXiv:2608. 16259v1 Announce Type: cross Abstract: The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable.
Beacon is a new agentic visual reasoning model that improves multimodal large language models (MLLMs) by better deciding when to use tools and how to use them. It introduces two key concepts—Mode Adaptiveness, which ensures tools are invoked only when necessary, and Tool Effect, which measures the net benefit of tool use— and trains the model with supervised fine‑tuning and reinforcement learning that rewards necessity-aware decisions and expands capability through expert hints. Across 13 benchmarks, Beacon outperforms other open‑source models, achieving the highest average score and the largest net tool‑gain on diagnostic tests.
arXiv:2607. 14256v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases.
The paper investigates an agentic framework for open‑world fake image detection that combines specialist detectors with per‑detector triage, prompting, and conflict‑aware evidence arbitration. Experiments across six configurations and three multimodal large language model backbones reveal that naive detector fusion yields high false‑positive rates, while triage and prompting consistently filter unreliable evidence. The most significant improvement comes from the reasoning component: a stronger judge markedly outperforms a weaker one, especially under distribution shift, and overall manipulation recall is nearly saturated, highlighting that the key challenge lies in calibrating trust and arbitrating conflicting forensic evidence rather than detecting manipulations themselves.
The paper introduces the Necessary Tool‑Evidence Path (NTEP) annotation scheme and its associated reward mechanism (NTEP‑R) to better supervise vision‑language models that use external tools. By explicitly specifying which evidence is needed and penalizing redundant tool calls, the authors train an 8B‑parameter model that shows improved accuracy and tool‑use efficiency across seven image‑grounded benchmarks. The approach demonstrates that fine‑grained supervision of tool‑evidence paths is essential for robust agentic VLM performance.