← Back to all news
arXiv Computer Vision August 24, 2026 By Chunyi Peng, Zhipeng Xu, Yukun Yan, Zhenghao Liu, Shi Yu, Sen Mei, Yubo Sun, Yongheng Zhang, Jie Zhou, Yu Gu, Ge Yu, Maosong Sun

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • rag
  • computer-vision
  • multimodal
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Aug 16

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across t...

llmsragcomputer-visionmultimodalbenchmarkssafety
More like this →
Hugging Face Trending Papers
Aug 2

VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval

Visual document retrieval has recently become increasingly important in applications such as enterprise search, scientific literature discovery, and retrieval-augmented generation. These applications depend on efficiently identifying query-relevant pages across large collections of visually rich documents.

ragbenchmarks
More like this →
arXiv AI
Jun 30

Multimodal Representation Alignment for Cross-modal Information Retrieval

arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.

By Fan Xu, Luis A. Leiva
llmsragmultimodalbenchmarkssafety
More like this →
arXiv Computer Vision
1d ago

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

arXiv:2608.21450v1 Announce Type: new Abstract: Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, e...

By Hangrui Xu, Zhengxian Wu, Yunyao Yu, Zhuohong Chen, Rui Cong, Xiangwen Deng, Zhifang Liu, Peng Jiao, Haoqian Wang
ragnlpefficiencysafety
More like this →
arXiv AI
Jun 10

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

arXiv:2606. 10572v1 Announce Type: new Abstract: External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence.

By Zhi Zheng, Ziqiao Meng, Hao Luan, Wei Liu, Wee Sun Lee
llmsragnlpefficiencymultimodalbenchmarks
More like this →
arXiv AI
Jul 24

KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback

arXiv:2607. 20556v1 Announce Type: new Abstract: In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis.

By Yan Zhu, Y. Chen, Rebecca Faust
llmsragsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea