CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.
By Dat Nguyen, Duc-Duy Nguyen
arXiv:2607. 18625v1 Announce Type: cross Abstract: Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones.
By Jin Yu, Juyoun Park
Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones. However, MambaOut demonstrates that a Gated CNN block can match or exceed VMamba on image classification, questioning the necessity of SSMs for vision.
The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.
By Soshun Kihara, Shunsuke Yasuki, Masato Taki
UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.
By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv:2609.21012v1 Announce Type: cross
Abstract: Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconogr...
By Sara Miketek, Biagio Barchielli, Nadeem Iqbal Kajla, Sinem Aslan
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
arXiv:2607. 25527v1 Announce Type: cross Abstract: Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities.
By Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li, Chen Chen, Sina Sajadmanesh, Lingjuan Lyu
DirectUV is an image-conditioned UV texture diffusion framework that generates high-quality textures directly in UV space using a pretrained image VAE and a Diffusion Transformer. It introduces Surface-Aware Positional Encoding (SAPE), which replaces standard 2D-grid positional encoding with encodings derived from 3D surface coordinates, enabling attention to operate based on surface proximity rather than UV-grid distance. A multi-level extension assigns different attention heads to progressively finer subdivisions of the UV patch, allowing the model to reason about surface structure at multiple granularities, resulting in sharper and more globally consistent textures, especially in occluded and view-unseen regions.
By Jiantao Lin, Yingjie Xu, Mingzhi Sheng, Yangkai Wei, Hao Chen, Ying-Cong Chen
arXiv:2609.00272v1 Announce Type: new
Abstract: Most advances in keypoint descriptions address monomodal settings, where image variations arise from viewpoint, illumination, or contrast changes. Mult...
By Paul Schneider, Nazim Haouchine
arXiv:2607.02486v2 Announce Type: replace
Abstract: Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet it...
By Yejun Zhang, Xinjue Wang, Zihan Wang, Esa Rahtu, Juho Kannala
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
By Kaustubh Kapil, Kishor P. Upla