Multi-Way Representation Alignment
arXiv:2602. 06205v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.
arXiv:2602. 06205v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.
The paper extends the analysis of Joint-Embedding Predictive Architectures (JEPAs) beyond Euclidean latent spaces to Riemannian manifolds. It shows that when latent variables lie on a sphere and the target distribution matches this spherical geometry, every optimal representation recovers the latent state up to an orthogonal transformation, demonstrating that Gaussian uniqueness is not universal. Experiments confirm that geometrically compatible targets improve linear recovery, especially in high-dimensional toroidal settings.
arXiv:2605. 30705v2 Announce Type: replace-cross Abstract: Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency.
arXiv:2601.21831v3 Announce Type: replace Abstract: We propose a geometric latent-subspace framework for generative modeling of discrete data. Specifically, we introduce latent subspaces in the expon...
arXiv:2602. 23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world.
arXiv:2606. 29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets.
Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models operate purely in a feed-forward, deterministic regime, \ie~they cannot generate plausible geometry beyond what the input views directly support.
arXiv:2607. 14228v1 Announce Type: cross Abstract: In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space.
arXiv:2603.02263v3 Announce Type: replace-cross Abstract: World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such...
arXiv:2511. 21594v3 Announce Type: replace Abstract: Large language models (LLMs) achieve state-of-the-art results across many natural language tasks, but their internal mechanisms remain difficult to interpret.
arXiv:2608. 16245v1 Announce Type: new Abstract: Disentangled representation learning seeks latent representations whose indicidual dimensions each align with a distinct covariate.
arXiv:2510. 09468v3 Announce Type: replace Abstract: Latent manifolds of autoencoders provide low-dimensional representations of data, which can be studied from a geometric perspective.