arXiv:2602. 24264v2 Announce Type: replace-cross Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems.
By Arnas Uselis, Andrea Dittadi, Seong Joon Oh
arXiv:2510.03075v4 Announce Type: replace-cross
Abstract: Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models....
By Karim Farid, Rajat Sahay, Yumna Ali Alnaggar, Simon Schrodi, Volker Fischer, Cordelia Schmid, Thomas Brox
arXiv:2608. 06809v1 Announce Type: new Abstract: How can an analyst decide whether a nonlinear dimensionality reduction embedding can be trusted?
By Xinyu Zhang, Klaus Mueller
arXiv:2606. 00124v1 Announce Type: cross Abstract: Positional embeddings (PEs) in Vision Transformers (ViTs) are known to impact performance and robustness, but their role in shaping internal spatial representations is not well understood.
By Mahmoud Mannes
arXiv:2603. 22278v2 Announce Type: replace-cross Abstract: Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations.
By Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba, Tamar Rott Shaham
arXiv:2609.05575v1 Announce Type: new
Abstract: Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability,...
By Yiming Tang, Harshvardhan Saini, Samyak Jha, Huaming Chen, Xufeng Duan, Dianbo Liu
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
By Kaustubh Kapil, Kishor P. Upla
arXiv:2607. 05568v1 Announce Type: cross Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding.
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
arXiv:2602.02611v2 Announce Type: replace
Abstract: A prevailing paradigm in modern representation learning is the map-first approach, in which a representation map is learned from reconstruction, em...
By David Vigouroux (ANITI, IMT Atlantique - DSD, LaTIM), Lucas Drumetz (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Ronan Fablet (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Fran\c{c}ois Rousseau (IMT Atlantique - DSD, LaTIM)
arXiv:2609.23717v1 Announce Type: new
Abstract: Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which regi...
By Liuyang Song, Yi Zhang, Zhongyi Deng, Daqian Yang, Hongbo Zhang
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can...
arXiv:2512. 08854v3 Announce Type: replace-cross Abstract: It has been hypothesized that achieving the data efficiency of human visual perception requires a generative approach in which internal representations result from inverting a decoder.
By Jack Brady, Bernhard Sch\"olkopf, Thomas Kipf, Simon Buchholz, Wieland Brendel