arXiv Computer Vision By Zuomin Qu

A Multi-View and Confusion-Guided Ensemble Framework for Robust Synthetic Image Attribution

Read the original on arXiv Computer Vision →

The paper introduces a multi‑view, confusion‑guided ensemble framework for synthetic image attribution, combining FFT‑ConvNeXt, DINOv2, CLIP, and Xception to capture frequency, semantic, and forensic cues. Extensive data augmentation simulates realistic post‑processing, while a binary expert classifier and class‑adaptive confidence calibration address ambiguities between similar diffusion models. The approach achieved 99.53% on the public leaderboard and 99.20% on the private leaderboard for the ICANN 2026 DLMMDD Workshop challenge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.