← Back to all news
arXiv Computer Vision August 31, 2026 By Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim

Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • rag
  • computer-vision
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 12

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

arXiv:2606. 13288v1 Announce Type: cross Abstract: Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding.

By Wei Li, Zhen Huang, Xinmei Tian
llmsdiffusionmultimodalbenchmarks
More like this →
arXiv Machine Learning
3d ago

SynMulti: Synthetic-to-Real Learning for Multimodal Video Understanding

arXiv:2604.12335v2 Announce Type: replace-cross Abstract: Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as...

By Tanzila Rahman, Renjie Liao, Leonid Sigal
llmscomputer-visionnlpfine-tuningmultimodal
More like this →
Hugging Face Trending Papers
Jun 2

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics

Vision-Language Models (VLMs) have demonstrated remarkable capabilities but suffer from significant computational overhead during inference. While visual token pruning offers a promising solution, existing methods predominantly rely on initial attention scores.

llmsefficiencymultimodalsafety
More like this →
arXiv AI
Jun 3

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics

arXiv:2606. 03569v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable capabilities but suffer from significant computational overhead during inference.

By Jiahui Wang, Kai Zhang, Mai Han, Huanghe Zhang
llmsefficiencymultimodalsafety
More like this →
arXiv AI
Jun 30

Steerable Visual Representations

arXiv:2604. 02327v2 Announce Type: replace-cross Abstract: Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation.

By Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi, Yuki M. Asano
llmscomputer-visionmultimodalbenchmarks
More like this →
arXiv AI
Jul 21

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

arXiv:2607. 17712v1 Announce Type: new Abstract: Detecting high-level semantic concepts like negation across modalities remains a challenge for current multimodal systems.

By Ali AbuSaleh, Leon Hammerla, Alexander Mehler
llmsragmultimodal
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea