The paper introduces the Real‑Calibrated Synthetic‑First Data Engine, a modular pipeline that integrates controllable diffusion‑based synthetic image generation with multi‑stage curation, filtering, and optional uncertainty‑driven selection and human verification. Designed as a CLI‑based framework, it allows independent configuration of generation, filtering, selection, and validation modules to enhance reproducibility and flexibility in real‑world data workflows. Empirical tests on human pose estimation demonstrate that synthetic data can boost a real‑data baseline when used as low‑cost augmentation, though synthetic‑only training still lags behind real‑only performance, underscoring the importance of data‑centric orchestration in low‑data regimes.
By Yukang Shen, Zhiguo Liu, Yingshu Li, Yan Huang
arXiv:2512.17730v2 Announce Type: replace
Abstract: Detectors of AI-generated images tend to inherit the biases of the data they are trained on: models fitted to GAN imagery learn to treat GAN-specif...
By Yichen Jiang, Mohammed Talha Alam, Sohail Ahmed Khan, Duc-Tien Dang-Nguyen, Fakhri Karray
arXiv:2605.21244v2 Announce Type: replace
Abstract: Super-Resolution (SR) has advanced rapidly in recent years, with diffusion-based models achieving unprecedented fidelity at the cost of introducing...
By Artem Borisov, Evgeney Bogatyrev, Khaled Abud, Dmitriy Vatolin
The paper introduces a multi‑view, confusion‑guided ensemble framework for synthetic image attribution, combining FFT‑ConvNeXt, DINOv2, CLIP, and Xception to capture frequency, semantic, and forensic cues. Extensive data augmentation simulates realistic post‑processing, while a binary expert classifier and class‑adaptive confidence calibration address ambiguities between similar diffusion models. The approach achieved 99.53% on the public leaderboard and 99.20% on the private leaderboard for the ICANN 2026 DLMMDD Workshop challenge.
By Zuomin Qu
arXiv:2607. 18230v1 Announce Type: cross Abstract: Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts.
By Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed Elhagry, Salwa K. Al Khatib, Tianjun Yao, Yonina C. Eldar, Jing-Hao Xue, Hao Li, Salman Khan, Zhiqiang Shen
arXiv:2607. 16283v1 Announce Type: cross Abstract: The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter generators they have not seen before.
By Md Faraz Kabir Khan, Saeed Anwar, Ghulam Mubashar Hassan
The paper proposes Artifact-Complementary Expert Fusion (ACEF), a two‑stage framework that enhances AI‑generated image detection by combining two types of reconstruction artifacts—VAE/DDIM and SRGAN—into aligned synthetic negatives. ACEF first builds artifact‑specific experts using LoRA adaptation on a frozen backbone, then fuses their multi‑layer evidence with Layer‑wise Artifact‑Complementary Fusion (LACF) to mitigate conflicts between artifact manifolds. Experiments on 13 benchmarks show that this approach improves generalizability over existing state‑of‑the‑art methods.
By Yiheng Li, Yang Yang, Wenhao Wang, Zichang Tan, Zecheng Lin, Li Gao, Zhen Lei
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model. Most recent methods adapt large pretrained text-to-image (T2I) latent diffusion models for their strong capacity and generative priors.
The paper introduces a multimodal foundation model for lunar remote sensing, trained from scratch on SomBench—a dataset of nearly two million co‑registered tile bundles across 11 modalities at 1 m and 100 m resolutions. The model extends the TerraMind masked‑token architecture with lunar‑specific features such as explicit acquisition geometry and joint training of two spatial scales, and employs FlexiViT patch embeddings for adaptable patch sizes. Evaluation on crater detection, irregular mare patch segmentation, and polar ice prospectivity regression shows that the pretrained model matches or surpasses ImageNet‑pretrained baselines, with notable label efficiency and effective adaptation via LoRA.
By Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal, Rachel A. Slank, Geoffrey Dawson, Hiyam Debary, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Nikolaos Dionelis, Ankur Kumar, Campbell D. Watson, Manil Maskey, Rebekah I. Dawson-Rigas, Juan Bernab\'e-Moreno, Rahul Ramachandran, Sujit Roy
arXiv:2606. 16082v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA).
By Guanyi Qin, Junjie Zhang, Chunming He, Yibing Fu, Jie Liang, Tianhe Wu, Lei Zhang
The paper introduces UnInfo, a test‑time adaptation method for vision‑language models like CLIP that addresses image corruption—a realistic distribution shift caused by sensor conditions. UnInfo leverages uniformity‑aware confidence maximization, information‑aware loss balancing, and knowledge distillation from an EMA teacher to preserve embedding uniformity and improve zero‑shot classification accuracy. Experiments show that UnInfo outperforms existing TTA methods on corrupted image datasets.
By Kazuki Adachi, Shin'ya Yamaguchi, Tomoki Hamagami
FoCLIP is a framework that creates adversarial examples to manipulate CLIP-based image quality metrics by reducing the alignment between image and text features. It uses stochastic gradient descent to combine feature alignment, score distribution balancing, and pixel‑guard regularization, enabling high CLIPscore predictions while maintaining visual fidelity. Experiments on artistic prompts and ImageNet show significant CLIPscore gains, and the authors also propose a color‑channel sensitivity detection method that achieves 91% accuracy.
By Yulin Chen, Zeyuan Wang, Tianyuan Yu, Yingmei Wei, Liang Bai