arXiv:2606. 23885v1 Announce Type: cross Abstract: Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder.
By Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi
arXiv:2608. 09997v1 Announce Type: new Abstract: Transformers have had a profound impact on the world of language processing and computer vision.
By Kaustubh Kapil, Kishor P. Upla
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
By Kaustubh Kapil, Kishor P. Upla
arXiv:2607. 29008v1 Announce Type: cross Abstract: Modern opaque AI models prize performance over interpretability, which makes testing difficult.
By Tyler Ashoff, Jordan Rodu
Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging.
arXiv:2606. 25293v1 Announce Type: new Abstract: Positional encodings (PEs) are essential for Transformers.
By Yipeng Zhang, Zhongtian Sun, Pietro Li\`o, Kelin Xia
arXiv:2606. 06342v1 Announce Type: cross Abstract: Topological Data Analysis (TDA) offers a principled, intrinsic lens for comparing neural representations.
By Yan Wang, Tianyang Hu
arXiv:2607. 19973v1 Announce Type: new Abstract: AI researchers describe state-of-the-art models as one thing repeated at scale: the Transformer, wired identically for text, pixels, or speech.
By Jaeho Seol
arXiv:2606. 09287v1 Announce Type: new Abstract: Understanding how transformer representations evolve across layers, not merely what they encode, remains an open problem in mechanistic interpretability.
By Vishal Pandey, Gopal Singh
arXiv:2606. 09770v1 Announce Type: cross Abstract: Nearby neurons in cortex share similar response profiles, producing systematic spatial organization across sensory and cognitive systems.
By Badr AlKhamissi, Johannes Mehrer, Lara Marinov, Ahmed Abdelaal, Abdulkadir Gokce, Martin Schrimpf
arXiv:2606. 19542v1 Announce Type: new Abstract: Large language models are commonly aligned through supervised fine-tuning, yet little is known about how their internal representations evolve during this process.
By Naman Malhotra, Jay Ambadkar, Abhinav Gupta, Kushal Kasivel, Abbas Schwarz, Kamillo Ferry, Anthea Monod
The growing number of medical vision foundation models highlights the need for effective model selection. However, mainstream selection methods rely on exhaustive fine-tuning, which is computationally expensive.