arXiv AI By Yusong Zhao, Hengyi Wang, Tanuja Ganu, Akshay Nambi, Hao Wang

Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs

Read the original on arXiv AI →

arXiv:2606. 16193v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 7

TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models

arXiv:2607. 04593v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model.

By Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda, Alaa Eddine Mazouz, Van-Tam Nguyen