Efficient Concertormer for Image Deblurring and Beyond
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2505.16157v3 Announce Type: replace Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Tran...
PixelUMM is an encoder‑free model that unifies image and video understanding and generation directly in pixel space. It represents images as spatial patches and videos as spatiotemporal tubelets, feeding both through single‑layer linear projections into a shared multimodal backbone. The Mixture‑of‑Transformers architecture blends shared attention with task‑specific parameters, enabling autoregressive text prediction, pixel‑space flow matching, and clean‑pixel video generation, and experiments show competitive performance across tasks while providing design insights for future pixel‑space multimodal models.
arXiv:2509.22650v3 Announce Type: replace Abstract: Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models,...
arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.
The paper introduces AttWarp, a lightweight technique that uses a multimodal large language model’s cross‑modal attention to perform rectilinear warping of input images at test time. By reallocating spatial resolution toward query‑relevant regions without altering model weights or architecture, AttWarp preserves global context while making small objects and subtle relationships easier for the model to read. Experiments on five benchmarks and four MLLMs show consistent accuracy gains, improved compositional reasoning, and reduced hallucinations compared to baseline image‑manipulation methods.
arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...