Fragment-Aware Vision Transformers for Fresco-Fragment Style Classification
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.
CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.
arXiv:2607.09086v2 Announce Type: replace Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
HiPerViT is a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based pipeline for texture recognition. It combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. Across six texture recognition benchmarks, HiPerViT consistently outperforms strong vision-only baselines, achieving notable gains on DTD, GTOS-Mobile, and 1200Tex, and the improvements are largely independent of backbone depth or fusion topology.
arXiv:2602.14633v3 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categori...
arXiv:2606. 09855v1 Announce Type: cross Abstract: Korean folk painting (minhwa) is built from a small vocabulary of auspicious symbols, a tiger for protection, a pair of birds for marital harmony, a peony for wealth, that recur across many of its painted genres.