arXiv:2605. 07914v2 Announce Type: replace Abstract: Sharpness-aware and gradient-alignment methods have been shown to improve generalization, however each family of methods targets a single geometric property of the loss landscape, while ignoring the other.
By Aristotelis Ballas, Christos Diou
arXiv:2602. 23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world.
By Simon Roschmann, Paul Krzakala, Sonia Mazelet, Quentin Bouniot, Zeynep Akata
arXiv:2609.38925v1 Announce Type: new
Abstract: Multimodal federated learning (MFL) has emerged as a pivotal paradigm for leveraging distributed data to enhance model performance. However, existing m...
By Tianchi Liao Tianchi_Liao, Lele Fu, Sheng Huang, Qing Hu, Hong-Ning Dai, Chuan Chen
arXiv:2606. 02172v1 Announce Type: new Abstract: Learning discriminative visual representations from distributed, heterogeneous data is a fundamental challenge in Federated Learning (FL).
By Mario Casado-Diez, Alejandro Dopico-Castro, Ver\'onica Bol\'on-Canedo, Bertha Guijarro-Berdi\~nas
arXiv:2606. 25347v1 Announce Type: new Abstract: Exemplar-free class-incremental learning (EFCIL) requires stable decision boundaries within a shifting feature space.
By Hongye Xu, Bartosz Krawczyk
arXiv:2608. 04234v1 Announce Type: cross Abstract: We study the problem of aligning data from multiple modalities into a shared representation space, focusing on settings where strong pretrained unimodal encoders are available but cross-modal paired data are scarce.
By Yixuan Florence Wu, Yilun Zhu, Naichen Shi
arXiv:2606. 16655v1 Announce Type: new Abstract: One-Shot Federated Learning (OSFL) addresses extreme communication regimes in which clients interact with the server only once, amplifying the impact of heterogeneous client data distributions.
By Daniele Berardini (AI for Good), Vito Paolo Pastore (AI for Good, MaLGa-DIBRIS, University of Genoa, Genoa, Italy), Vittorio Murino (AI for Good, Department of Computer Science, University of Verona, Verona, Italy)
LLaVAFlow is an information‑theoretic distillation framework designed to preserve cross‑modal alignment in Multimodal Large Language Models during visual instruction tuning. It compresses the mutual information between extracted relations and MLLM embeddings to refine alignment flow, and then maximizes mutual information between pretrained and fine‑tuned alignment flows to transfer compact alignment information. Experiments demonstrate that LLaVAFlow effectively maintains alignment flow, improving downstream performance and generalization.
By Muyao Yuan, Muyan Jiao, Jiangyong Ying, Weizhan Zhang, Yuanhong Zhang, Lan Ma, Yuan Gao, Haipeng Du
DMM-Align introduces a closed‑loop framework for 2D‑3D registration that jointly refines correspondences, estimates pose, and learns representations using a shared differentiable geometric state. The method employs two diffusion processes: a geometry‑aware diffusion that improves the soft matching matrix for robust correspondence estimation, and a geometry‑conditioned diffusion teacher that feeds pose‑induced supervision back into feature learning. Experiments on 7‑Scenes and RGB‑D Scenes V2 show that DMM‑Align outperforms strong baselines, particularly in low‑overlap and heavily occluded scenarios, demonstrating the value of closed‑loop geometric feedback.
By Chongjian Wang, Junjie Gao
arXiv:2605.13155v2 Announce Type: replace
Abstract: Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward...
By Ying Ba, Tianyu Zhang, Mohan Zhou, Yalong Bai, Wenyi Mo, Guiwei Zhang, Bing Su, Ji-Rong Wen
arXiv:2604. 27147v3 Announce Type: replace-cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human preferences, a problem known as \textit{guidance}.
By Jerry Y. Huang, Justin Lin, Sheel Shah, Kartik Nair, Nicholas M. Boffi
Spectral Feedback is a new algorithm for aligning discrete diffusion models at test time by iteratively revisiting and editing token positions rather than only steering the reverse process. It selects edit-sets—groups of token positions to re-mask and re-sample—using sparse Fourier representations of edit-set value functions, enabling efficient optimization of which tokens to revisit. The method is model-agnostic and improves alignment performance across pretrained, test‑time aligned, and fine‑tuned diffusion models, achieving significant gains in protein stability for inverse folding tasks.
By Shai Dickman, Mert Cemri, Landon Butler, Kannan Ramchandran