Team MSU GenText-Forensics Challenge 2026 Technical Report
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces an evidence‑guided detector‑localizer‑reasoner system for text‑centric image forensics, addressing the challenges posed by AI‑generated content. It combines an image‑level authenticity detector, a localizer that extracts tampered regions, and an MLLM‑based reasoner that generates structured forensic reports grounded in the detected evidence. The system employs iterative difficulty‑aware mining and report‑mask consistency post‑processing, achieving a score of 0.638 and ranking second in the ACM Multimedia 2026 GenText‑Forensics Challenge.
arXiv:2609.36145v1 Announce Type: new Abstract: Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective...
ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.
arXiv:2609.39066v1 Announce Type: new Abstract: Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large l...
FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.
arXiv:2605. 09089v2 Announce Type: replace-cross Abstract: Digital onboarding and eKYC systems used by banks, fintech platforms, telecom providers, and other third-party services commonly verify users by comparing an uploaded identity document with a selfie or live facial capture.