Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,374 stories · RSS feed

arXiv Computer Vision
Oct 2

SR-Ground: Image Quality Grounding for Super-Resolved Content

SR‑Ground is a large‑scale dataset created to enable fine‑grained segmentation of visual artifacts in super‑resolved images. It contains 63,000 images processed by various state‑of‑the‑art SR models, each annotated at the pixel level for six distinct artifact types, validated through a crowdsourcing study with 1,062 participants. The dataset improves the training of image quality assessment models with grounding capabilities and supports a fine‑tuning pipeline that reduces perceptible artifacts in SR outputs, outperforming no‑reference methods on both benchmark and real‑world low‑resolution datasets.

By Artem Borisov, Evgeney Bogatyrev, Khaled Abud, Dmitriy Vatolin
arXiv AI
Oct 2

DeFA: Dependency-Guided Failure Attribution for LLM Agents

DeFA is a dependency-guided framework that attributes failures in large language model agents by constructing an event dependency graph and a failure propagation graph from protocol relations and semantic dependencies. It identifies violating events, traces their sources and effects, and determines the decisive error, responsible agent, and error category. The method supports long trajectories through segmentation and has shown superior accuracy on text, image, and video tasks, while its diagnostic feedback can improve agent performance on subsequent tasks.

By Bo Deng, Xinlei Zheng, Yi Wei, Kang Zhou, Chongyang Tao, Renzhao Liang, Xuanren Chen, Lifan Guo, Chi Zhang
arXiv AI
Oct 2

Contrastive Attention Mitigates Spectral Bias in Spiking Transformers

The paper introduces Spiking Contrastive Attention (SCA), a module designed to reduce spectral bias in Spiking Transformers by enhancing high‑frequency information. It demonstrates that spiking neurons and spiking self‑attention act as low‑pass filters, leading to loss of high‑frequency components. Experiments show that SCA improves performance across image classification, semantic segmentation, and event‑based tracking while maintaining lower complexity than the original spiking self‑attention.

By Xiaoli Liu, Malu Zhang, Yang Yang
arXiv AI
Oct 2

SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials

The paper introduces SW-KAN, a Kolmogorov‑Arnold Network that replaces traditional B‑spline activations with Stieltjes‑Wigert q‑orthogonal polynomials defined on the semi‑infinite domain (0, ∞). It addresses the domain mismatch between unbounded inputs and bounded polynomial bases by applying a smooth exponential‑of‑tanh mapping, and uses a numerically stable three‑term recurrence to evaluate polynomial expansions efficiently. Experiments on image classification and continuous function approximation show that SW‑KAN achieves better accuracy‑efficiency trade‑offs than existing polynomial KANs, especially in resource‑constrained scenarios with limited data or feature dimensionality.

By Amirhosein Azarpour, Seyyed Moein Kazemi
arXiv AI
Oct 2

VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision

VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.

By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
arXiv Computer Vision
Oct 2

PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion

PixelDense introduces a dual‑stream representation alignment for pixel diffusion, separating semantic and geometric teachers (DINOv2, SAM2, Depth Anything v2, Metric3D v2) into distinct projection spaces with an orthogonality penalty. The method improves dense‑prediction benchmarks, boosting PixelGen‑XXL’s GenEval score from 0.7927 to 0.8093, achieving significant gains in panoptic quality and depth accuracy, and accelerating training from random initialization. It also enhances SDEdit editing by preserving background structure and increasing PSNR.

By Lehan Yang, Daiqing Qi, Wenhao Zhang, Avery Li, Yiqing Yang, Yifan Li, Yu Kong, Haitian Zheng, Zhifei Zhang, Zhe Lin, Varun Jampani, Sheng Li
arXiv AI
Oct 2

RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation

The paper introduces RelationVGGT, a feed‑forward framework that performs 3D spatial relation segmentation without per‑scene optimization or known camera poses. It combines semantic features from a visual foundation model with geometry‑aware representations from a 3D geometry foundation model, and uses a relation transformer to predict subject‑conditioned, cross‑view relations based on a visual subject and a textual query. The authors also present an automated annotation pipeline built on ScanNet++ with VLMs and LLMs to generate scalable training data for this new task.

By Minsu Kim, Jaesung Choe, Jiwoo Lee, Yu-Chiang Frank Wang, Seon Joo Kim
arXiv AI
Oct 2

CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

CineMR is a vision‑language model that integrates cardiac image‑analysis tools to perform quantitative assessment of cine cardiac MRI. It uses supervised fine‑tuning and Group Relative Policy Optimization to learn reliable tool invocation, achieving significantly higher accuracy on a multi‑cohort benchmark than existing medical VLMs. The model demonstrates that tool‑augmented reasoning improves ventricular measurement accuracy by up to 23.7% and is essential for robust quantitative CMR interpretation.

By Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang
arXiv Computer Vision
Oct 2

EgoRefine: Ego-Referenced Predictive Alignment and Trajectory-Conditioned Reliability-Aware Fusion for Asynchronous Collaborative Perception

arXiv:2610.00319v1 Announce Type: new Abstract: Collaborative perception enables connected agents to share complementary observations for 3D object detection, extending sensing range and mitigating o...

By Lingzhao Kong, Yongsheng Zang, Yu Kang, Kailun Yang, Jie Fu, Yukun Zuo, Zhiyong Li
arXiv Computer Vision
Oct 2

Vmem-$\varphi$: Low-Compute Out-of-Distribution Detection in Spiking Neural Networks from Membrane-Potential Statistics

arXiv:2610.00350v1 Announce Type: new Abstract: Spiking Neural Networks (SNNs) offer an energy-efficient approach to processing event-camera data, yet out-of-distribution (OOD) detection remains chal...

By Arul Rana, Agrim Tripathi, Shoaib Ahmed Dipu, Md. Shaown Miah, Syed Ishtiaque Ahmed, Sayeed Shafayet Chowdhury
arXiv Computer Vision
Oct 2

Machine Translation for Sign Languages

arXiv:2610.00881v1 Announce Type: new Abstract: Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation...

By Ozge Mercanoglu Sincan, Anton Pelykh, Edward Fish, Harry Walsh, JianHe Low, Karahan Sahin, Oline Ranum, Sobhan Asasi, Steven Emery, Richard Bowden