Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.36458v1 Announce Type: new Abstract: Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions...
arXiv:2602.02611v2 Announce Type: replace Abstract: A prevailing paradigm in modern representation learning is the map-first approach, in which a representation map is learned from reconstruction, em...
arXiv:2606. 00124v1 Announce Type: cross Abstract: Positional embeddings (PEs) in Vision Transformers (ViTs) are known to impact performance and robustness, but their role in shaping internal spatial representations is not well understood.
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current models use this scale inf...
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
arXiv:2609.38285v1 Announce Type: cross Abstract: Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Add...