← Back to all news
arXiv Computer Vision August 25, 2026 By Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang

ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • fine-tuning
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
2d ago

ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in light...

llmsfine-tuningmultimodalbenchmarks
More like this →
arXiv AI
Aug 10

RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

arXiv:2608. 07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive.

By Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo
llmsefficiencymultimodal
More like this →
arXiv AI
Jun 2

On the Limits of Token Reduction for Efficient Unified Vision Language Training

arXiv:2606. 01503v1 Announce Type: cross Abstract: Unified vision-language models (VLMs) integrate visual understanding and visual generation within a single autoregressive backbone, but their joint training is computationally expensive and largely overlooked from an efficiency perspective.

By Siyi Chen, Weiming Zhuang, Jingtao Li, Lingjuan Lv
llmsmultimodal
More like this →
arXiv AI
Aug 18

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

arXiv:2509. 06461v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable success across diverse visual tasks, yet their performance degrades in complex visual environments.

By Yuyao Ge, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Xuanshan Zhou, Jiayu Yao, Jiafeng Guo, Xueqi Cheng
llmscomputer-visionmultimodal
More like this →
arXiv AI
Jun 24

Listening makes Vision Clear for VLMs

arXiv:2606. 23763v1 Announce Type: cross Abstract: Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens.

By Yiyang Chen, Yixin Tan, Binrui Shen
safety
More like this →
arXiv Machine Learning
Jun 2

Self-Improving Small Object Grounding in LVLMs

arXiv:2606. 01612v1 Announce Type: cross Abstract: Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning?

By Tianze Yang, Yucheng Shi, Ruitong Sun, Ninghao Liu, Jin Sun
llmsfine-tuningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea