SR‑Ground is a large‑scale dataset created to enable fine‑grained segmentation of visual artifacts in super‑resolved images. It contains 63,000 images processed by various state‑of‑the‑art SR models, each annotated at the pixel level for six distinct artifact types, validated through a crowdsourcing study with 1,062 participants. The dataset improves the training of image quality assessment models with grounding capabilities and supports a fine‑tuning pipeline that reduces perceptible artifacts in SR outputs, outperforming no‑reference methods on both benchmark and real‑world low‑resolution datasets.
By Artem Borisov, Evgeney Bogatyrev, Khaled Abud, Dmitriy Vatolin
DeFA is a dependency-guided framework that attributes failures in large language model agents by constructing an event dependency graph and a failure propagation graph from protocol relations and semantic dependencies. It identifies violating events, traces their sources and effects, and determines the decisive error, responsible agent, and error category. The method supports long trajectories through segmentation and has shown superior accuracy on text, image, and video tasks, while its diagnostic feedback can improve agent performance on subsequent tasks.
By Bo Deng, Xinlei Zheng, Yi Wei, Kang Zhou, Chongyang Tao, Renzhao Liang, Xuanren Chen, Lifan Guo, Chi Zhang
The paper introduces Spiking Contrastive Attention (SCA), a module designed to reduce spectral bias in Spiking Transformers by enhancing high‑frequency information. It demonstrates that spiking neurons and spiking self‑attention act as low‑pass filters, leading to loss of high‑frequency components. Experiments show that SCA improves performance across image classification, semantic segmentation, and event‑based tracking while maintaining lower complexity than the original spiking self‑attention.
By Xiaoli Liu, Malu Zhang, Yang Yang
The paper introduces SW-KAN, a Kolmogorov‑Arnold Network that replaces traditional B‑spline activations with Stieltjes‑Wigert q‑orthogonal polynomials defined on the semi‑infinite domain (0, ∞). It addresses the domain mismatch between unbounded inputs and bounded polynomial bases by applying a smooth exponential‑of‑tanh mapping, and uses a numerically stable three‑term recurrence to evaluate polynomial expansions efficiently. Experiments on image classification and continuous function approximation show that SW‑KAN achieves better accuracy‑efficiency trade‑offs than existing polynomial KANs, especially in resource‑constrained scenarios with limited data or feature dimensionality.
By Amirhosein Azarpour, Seyyed Moein Kazemi
VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.
By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
PixelDense introduces a dual‑stream representation alignment for pixel diffusion, separating semantic and geometric teachers (DINOv2, SAM2, Depth Anything v2, Metric3D v2) into distinct projection spaces with an orthogonality penalty. The method improves dense‑prediction benchmarks, boosting PixelGen‑XXL’s GenEval score from 0.7927 to 0.8093, achieving significant gains in panoptic quality and depth accuracy, and accelerating training from random initialization. It also enhances SDEdit editing by preserving background structure and increasing PSNR.
By Lehan Yang, Daiqing Qi, Wenhao Zhang, Avery Li, Yiqing Yang, Yifan Li, Yu Kong, Haitian Zheng, Zhifei Zhang, Zhe Lin, Varun Jampani, Sheng Li
The paper introduces RelationVGGT, a feed‑forward framework that performs 3D spatial relation segmentation without per‑scene optimization or known camera poses. It combines semantic features from a visual foundation model with geometry‑aware representations from a 3D geometry foundation model, and uses a relation transformer to predict subject‑conditioned, cross‑view relations based on a visual subject and a textual query. The authors also present an automated annotation pipeline built on ScanNet++ with VLMs and LLMs to generate scalable training data for this new task.
By Minsu Kim, Jaesung Choe, Jiwoo Lee, Yu-Chiang Frank Wang, Seon Joo Kim
CineMR is a vision‑language model that integrates cardiac image‑analysis tools to perform quantitative assessment of cine cardiac MRI. It uses supervised fine‑tuning and Group Relative Policy Optimization to learn reliable tool invocation, achieving significantly higher accuracy on a multi‑cohort benchmark than existing medical VLMs. The model demonstrates that tool‑augmented reasoning improves ventricular measurement accuracy by up to 23.7% and is essential for robust quantitative CMR interpretation.
By Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang
arXiv:2601.09879v2 Announce Type: replace-cross
Abstract: Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report gen...
By Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
arXiv:2609.19011v2 Announce Type: replace
Abstract: Knowledge distillation can copy a deployed model by training a student on its logits or features. The student inherits a watermark only through the...
By Redwanul Karim, Tobias Feigl, Christopher Mutschler, Felix Ott
arXiv:2610.00040v1 Announce Type: new
Abstract: Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation mo...
By Thanh-Khoi Nguyen, Thien-Phuc Tran, Minh-Triet Tran
arXiv:2610.00319v1 Announce Type: new
Abstract: Collaborative perception enables connected agents to share complementary observations for 3D object detection, extending sensing range and mitigating o...
By Lingzhao Kong, Yongsheng Zang, Yu Kang, Kailun Yang, Jie Fu, Yukun Zuo, Zhiyong Li
arXiv:2610.00350v1 Announce Type: new
Abstract: Spiking Neural Networks (SNNs) offer an energy-efficient approach to processing event-camera data, yet out-of-distribution (OOD) detection remains chal...
By Arul Rana, Agrim Tripathi, Shoaib Ahmed Dipu, Md. Shaown Miah, Syed Ishtiaque Ahmed, Sayeed Shafayet Chowdhury
arXiv:2610.00881v1 Announce Type: new
Abstract: Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation...
By Ozge Mercanoglu Sincan, Anton Pelykh, Edward Fish, Harry Walsh, JianHe Low, Karahan Sahin, Oline Ranum, Sobhan Asasi, Steven Emery, Richard Bowden
arXiv:2610.01013v1 Announce Type: new
Abstract: Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pas...
By Junyi Wu, Fanqing Kong, Leyang Chen, Shaoqiu Zhang, Yulun Zhang
arXiv:2610.01022v1 Announce Type: new
Abstract: Memory-attention-based Video Instance Segmentation (VIS) methods have demonstrated strong zero-shot tracking capability, yet their substantial memory r...
By Arash Rocky, Q. M. Jonathan Wu
arXiv:2610.01409v1 Announce Type: new
Abstract: Reliable uncertainty estimation is essential for deploying object detectors when distribution/covariate shift and adversarial attacks may occur. Existi...
By Charmaine Barker, Daniel Bethell, Simos Gerasimou
arXiv:2610.01452v1 Announce Type: new
Abstract: While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastroph...
By Samuel Hart, Ahmad Yahya, Ahmed Karam Eldaly
arXiv:2610.01794v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that ot...
By Edward W. Staley, Connor O. Pyles, Rahul Hingorani, Frank Camargo, Griffin Milsap, Jared Markowitz, Matthew S. Fifer, Michael Wolmetz
arXiv:2610.01870v1 Announce Type: new
Abstract: Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventi...
By Hosam Elgendy, Utkarsh Mall