arXiv Computer Vision By Haiming Zhao, Tai Wang, Kun Zhang, Xicheng Peng, Zhiyang Li

Concept Driven Domain Adaptation: Finding an Abstract Needle in a Haystack

Read the original on arXiv Computer Vision →

The paper introduces Concept-Driven Domain Adaptation (CDDA), a three-stage framework that adapts vision‑language models for concept‑to‑example video retrieval in educational settings. CDDA first structures textual embeddings using textbook and teacher‑handbook concept pairs, then transfers this geometry to documentary visuals with a frozen visual encoder, and finally jointly fine‑tunes both encoders with sparse visual concept supervision. On a middle‑school physics benchmark, CDDA outperforms several multimodal baselines in retrieving concept‑driven moments while preserving concrete image‑text alignment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 24

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

arXiv:2608.15698v2 Announce Type: replace Abstract: Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from docume...

By Chunyi Peng, Zhipeng Xu, Yukun Yan, Zhenghao Liu, Shi Yu, Sen Mei, Yubo Sun, Yongheng Zhang, Jie Zhou, Yu Gu, Ge Yu, Maosong Sun
arXiv Computation and Language
Sep 4

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

KnowVis is a framework that converts linear video lectures into knowledge‑centric visual narratives. It first extracts a detailed concept map from multimodal video content to identify key and challenging concepts, then builds structured knowledge units and synthesizes engaging visual summaries. The authors also provide a curated dataset of 125 educational videos across 10 disciplines, paired with 1,079 visual summaries, and show through automated evaluations and a human study that KnowVis produces more accurate, clear visuals that reduce cognitive load and improve learning effectiveness and knowledge retention.

By Yi Xu, Yifan Hou, Xiaoyu Zhang