arXiv:2504. 18455v2 Announce Type: replace-cross Abstract: We study distributed multiview representation learning, a problem in which $K$ clients each observe a distinct but possibly statistically correlated view.
By Milad Sefidgaran, Piotr Krasnowski, Abdellatif Zaidi
The study investigates whether transferring relational structure from human mental representations to deep neural networks (DNNs) can improve fine‑grained alignment between the two. Using unsupervised Gromov‑Wasserstein optimal transport, the authors show that fine‑tuning pre‑trained DNNs with Relational Knowledge Distillation (RKD) brings the networks close enough to human representations to align at the individual‑object level on a test set of concepts not seen during training. The improvement is driven mainly by a more human‑like global structure of category distances, while local nearest‑neighbor overlap remains largely unchanged.
By Yuria Shimizu, Soh Takahashi, Takato Horii, Masafumi Oizumi
CLIP-RD introduces a relational distillation framework for efficient CLIP knowledge distillation, featuring Vertical Relational Distillation (VRD) and Cross Relational Distillation (XRD). VRD aligns intra‑modal similarity distributions between teacher and student, while XRD aligns cross‑modal similarity distributions to enforce bidirectional symmetry. This joint modeling of multidirectional relational structures improves the student’s embedding geometry, yielding a 1.8%p performance gain over CLIP‑KD across various architectures, tasks, and corruption settings with minimal training‑time overhead.
By Jeannie Chung, Hanna Jang, Ingyeong Yang, Uiwon Hwang, Jaehyeong Sim
The paper introduces SIMPLE, a prior‑fitted multi‑view in‑context learner that learns a reusable, task‑conditioned inference procedure instead of a fixed fusion function. By generating synthetic task priors in embedding space, SIMPLE can handle diverse view configurations, class structures, and missingness patterns. Experiments on multi‑view and multi‑omics benchmarks show that a frozen SIMPLE model performs competitively, and lightweight adapter calibration further improves performance across most datasets.
By Jielong Lu, Zhihao Wu, Jiajun Yu, Zhaoliang Chen, Haishuai Wang
Spatial-OPSD is a label‑free self‑improvement framework for vision‑language models that leverages spatial priors such as depth, 3D relations, and camera geometry to provide dense token‑level supervision. During training, a privileged teacher uses these priors while the student learns from only the original visual‑language input, and a recursive round‑wise scheme allows repeated self‑improvement without moving the teacher. Across four VLM families, one round of Spatial‑OPSD improves the five‑benchmark average, and three rounds push a strong spatially specialized model to the open‑source frontier, achieving the highest average among open models and best results on three of five spatial reasoning benchmarks.
By Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An, Haodong Jing, Ruqi Huang
arXiv:2606. 25927v1 Announce Type: cross Abstract: As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding.
By Luyang Fang, Haoran Lu, Yongkai Chen, Wenxuan Zhong, Ping Ma
arXiv:2608. 08309v1 Announce Type: cross Abstract: We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy.
By Nikos Giakoumoglou, Paschalis Giakoumoglou, Tania Stathaki
arXiv:2606. 00771v1 Announce Type: cross Abstract: A simple way to improve the performance of almost any machine learning model is not to train a single but several models with diverse algorithms which will make slightly distinct kinds of predictions and errors on the same data, and thus improve the average predictions and robustness.
By Yiru Yang, Junling Wang, Nishant Kumar Singh, Luohong Wu, Haoran Yan
arXiv:2608. 01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories.
By Leyan Xue, Feng Xiong, Mingjun Ma, Changqing Zhang
arXiv:2606. 29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets.
By Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon
arXiv:2608. 04234v1 Announce Type: cross Abstract: We study the problem of aligning data from multiple modalities into a shared representation space, focusing on settings where strong pretrained unimodal encoders are available but cross-modal paired data are scarce.
By Yixuan Florence Wu, Yilun Zhu, Naichen Shi
The paper proposes Self‑Distillation Fine‑Tuning (SDFT) as a method to recover performance in Large Language Models that has been degraded by catastrophic forgetting, quantization, or pruning. It shows that SDFT restores model capabilities by aligning the high‑dimensional manifold of the student model’s hidden layers with that of a teacher model, as measured by Centered Kernel Alignment (CKA). The authors provide both empirical evidence of strong correlation between manifold alignment and performance recovery and a theoretical explanation linking generative capability to the structure of these manifolds.
By Chi Liu, Xin Chen, Xu Zhou, Fangbo Tu, Srinivasan Manoharan