← Back to all news
arXiv Computation and Language August 24, 2026 By Pume Tuchinda, Parinthapat Pengpun, Romrawin Chumpu, Patomporn Payoungkhamdee, Sarana Nutanong, Peerat Limkonchotiwat

When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • nlp
  • efficiency
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Computer Vision
2d ago

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

arXiv:2608.25575v1 Announce Type: new Abstract: Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and rel...

By Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji
llmsdiffusionefficiencymultimodalbenchmarks
More like this →
Hugging Face Trending Papers
3d ago

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this...

llmsdiffusionefficiencymultimodalbenchmarks
More like this →
arXiv AI
Jun 29

Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge

arXiv:2606. 27527v1 Announce Type: cross Abstract: Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pretraining, yet their potential to supervise models in other modalities remains underexplored.

By Thomas Shih-Chao Liang, Zhuoran Yu, Yong Jae Lee
llmsefficiencymultimodalbenchmarks
More like this →
arXiv AI
Jun 3

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement

arXiv:2412. 01282v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks.

By Qianhan Feng, Wenshuo Li, Tong Lin, Xinghao Chen
llmsragefficiencymultimodalbenchmarkssafety
More like this →
arXiv Machine Learning
Aug 4

Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models

arXiv:2608. 01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories.

By Leyan Xue, Feng Xiong, Mingjun Ma, Changqing Zhang
llmsefficiencymultimodalbenchmarks
More like this →
arXiv AI
Aug 6

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

arXiv:2608. 05131v1 Announce Type: cross Abstract: On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs).

By Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
llmsefficiencymultimodalbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea