Equivariant Representation Learning via Class-Pose Decomposition
arXiv:2207. 03116v4 Announce Type: replace Abstract: We introduce a general method for learning representations that are equivariant to symmetries of data.
arXiv:2605. 30705v2 Announce Type: replace-cross Abstract: Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency.
arXiv:2207. 03116v4 Announce Type: replace Abstract: We introduce a general method for learning representations that are equivariant to symmetries of data.
arXiv:2606. 04108v1 Announce Type: cross Abstract: Single-view 3D generative models have achieved impressive visual quality, yet they are not designed to satisfy structural or functional requirements, and in practice, often fall short.
arXiv:2608.30388v1 Announce Type: cross Abstract: Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentri...
arXiv:2601.21831v3 Announce Type: replace Abstract: We propose a geometric latent-subspace framework for generative modeling of discrete data. Specifically, we introduce latent subspaces in the expon...
arXiv:2606. 26535v1 Announce Type: cross Abstract: Current VLM evaluations often conflate language priors with genuine spatial reasoning.
arXiv:2605.06140v3 Announce Type: replace-cross Abstract: Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such...
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning.
arXiv:2607. 00556v1 Announce Type: cross Abstract: While recent advancements like the Poincar\'e ResNet have demonstrated the potential of learning visual representations directly in hyperbolic space, their optimisation remains hampered by the computationally intensive nature of Riemannian gradients and the strict boundaries of the manifold.
arXiv:2608.28840v1 Announce Type: new Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries. These latent geometries are not directly compatible,...
arXiv:2505. 04486v4 Announce Type: replace-cross Abstract: Flow matching models have shown great potential in image generation tasks among probabilistic generative models.
arXiv:2603.02263v3 Announce Type: replace-cross Abstract: World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such...
arXiv:2602.02611v2 Announce Type: replace Abstract: A prevailing paradigm in modern representation learning is the map-first approach, in which a representation map is learned from reconstruction, em...