arXiv:2506.11784v2 Announce Type: replace
Abstract: Vision Transformers (ViTs) are essential in computer vision but are computationally intensive, too. Model quantization, particularly to low bit-wid...
By Guang Liang, Xinyao Liu, Jianxin Wu
arXiv:2510. 06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data.
By Ayush Zenith, Arnold Zumbrun, Neel Raut, Jing Lin
arXiv:2607. 16283v1 Announce Type: cross Abstract: The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter generators they have not seen before.
By Md Faraz Kabir Khan, Saeed Anwar, Ghulam Mubashar Hassan
arXiv:2607. 28589v1 Announce Type: cross Abstract: Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices.
By Md. Mehrab Hossain Opi, Robiul Islam Ryad, Md. Umar Faruk
VQ-Transplant is a framework that allows new vector‑quantization (VQ) modules to be inserted into frozen, pre‑trained visual tokenizers without retraining the entire model. By preserving all encoder‑decoder parameters and adding a lightweight decoder adaptation trained for only five epochs on ImageNet‑1k, the method mitigates decoder‑quantization mismatch. Experiments show that VQ-Transplant achieves near state‑of‑the‑art reconstruction fidelity for industry‑level models such as VAR while cutting training costs by 95%.
By Xianghong Fang, Yuan Yuan, Dehan Kong, Tim G. J. Rudner
arXiv:2605. 29539v2 Announce Type: replace-cross Abstract: Vision-language foundation models have shown promising zero-shot generalization for Cross-Domain Few-Shot Object Detection (CD-FSOD).
By Jiacong Liu, Shu Luo, Yikai Qin, Yaze Zhao, Yongwei Jiang, Yixiong Zou
arXiv:2609.16689v1 Announce Type: new
Abstract: Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained...
By Jinwoo Jeon, GyuYeop Do, Yubin Lim, Nam-Joon Kim, Hyun Gon Ryu, Hyuk-Jae Lee, Byung-Jun Lee
arXiv:2606. 30528v1 Announce Type: cross Abstract: Current generative models, including GANs and diffusion models, have reached an outstanding level of photorealism, posing significant risks to privacy and security.
By Orazio Pontorno, Mattia Litrico, Luca Guarnera, Mario Valerio Giuffrida, Sebastiano Battiato
Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained edge devices remains challenging. EdgeVL address...
OD3 introduces an optimization‑free dataset distillation framework tailored for object detection. The method first iteratively places object instances in synthesized images, then screens candidates with a pre‑trained observer model to discard low‑confidence objects. Applied to MS COCO and PASCAL VOC, OD3 achieves compression ratios from 0.25% to 5% and surpasses previous detection‑focused distillation methods by over 14% on COCO mAP50 at a 1.0% compression ratio.
By Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen
arXiv:2608. 03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped.
By Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong
arXiv:2607. 18230v1 Announce Type: cross Abstract: Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts.
By Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed Elhagry, Salwa K. Al Khatib, Tianjun Yao, Yonina C. Eldar, Jing-Hao Xue, Hao Li, Salman Khan, Zhiqiang Shen