arXiv Computer Vision By Jo\~ao Pedro C. A. de S\'a, Odemir Martinez Bruno

HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture Recognition

Read the original on arXiv Computer Vision →

HiPerViT is a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based pipeline for texture recognition. It combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. Across six texture recognition benchmarks, HiPerViT consistently outperforms strong vision-only baselines, achieving notable gains on DTD, GTOS-Mobile, and 1200Tex, and the improvements are largely independent of backbone depth or fusion topology.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 24

CRISP: Compositional Relations as Invariant Structural Priors for Domain Generalization

CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.

By Dat Nguyen, Duc-Duy Nguyen
arXiv Computer Vision
6d ago

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.

By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv Computer Vision
Sep 24

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.

By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun