arXiv Computer Vision
Aug 31

Relational Knowledge Distillation Brings DNN Representations Close Enough to Humans to Be Aligned Without Supervision

The study investigates whether transferring relational structure from human mental representations to deep neural networks (DNNs) can improve fine‑grained alignment between the two. Using unsupervised Gromov‑Wasserstein optimal transport, the authors show that fine‑tuning pre‑trained DNNs with Relational Knowledge Distillation (RKD) brings the networks close enough to human representations to align at the individual‑object level on a test set of concepts not seen during training. The improvement is driven mainly by a more human‑like global structure of category distances, while local nearest‑neighbor overlap remains largely unchanged.

By Yuria Shimizu, Soh Takahashi, Takato Horii, Masafumi Oizumi
arXiv Computer Vision
Sep 11

CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation

CLIP-RD introduces a relational distillation framework for efficient CLIP knowledge distillation, featuring Vertical Relational Distillation (VRD) and Cross Relational Distillation (XRD). VRD aligns intra‑modal similarity distributions between teacher and student, while XRD aligns cross‑modal similarity distributions to enforce bidirectional symmetry. This joint modeling of multidirectional relational structures improves the student’s embedding geometry, yielding a 1.8%p performance gain over CLIP‑KD across various architectures, tasks, and corruption settings with minimal training‑time overhead.

By Jeannie Chung, Hanna Jang, Ingyeong Yang, Uiwon Hwang, Jaehyeong Sim
arXiv Machine Learning
Aug 20

Pretraining Reusable Inference Across Views with Synthetic Task Priors

The paper introduces SIMPLE, a prior‑fitted multi‑view in‑context learner that learns a reusable, task‑conditioned inference procedure instead of a fixed fusion function. By generating synthetic task priors in embedding space, SIMPLE can handle diverse view configurations, class structures, and missingness patterns. Experiments on multi‑view and multi‑omics benchmarks show that a frozen SIMPLE model performs competitively, and lightweight adapter calibration further improves performance across most datasets.

By Jielong Lu, Zhihao Wu, Jiajun Yu, Zhaoliang Chen, Haishuai Wang
arXiv Computer Vision
4d ago

Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

Spatial-OPSD is a label‑free self‑improvement framework for vision‑language models that leverages spatial priors such as depth, 3D relations, and camera geometry to provide dense token‑level supervision. During training, a privileged teacher uses these priors while the student learns from only the original visual‑language input, and a recursive round‑wise scheme allows repeated self‑improvement without moving the teacher. Across four VLM families, one round of Spatial‑OPSD improves the five‑benchmark average, and three rounds push a strong spatially specialized model to the open‑source frontier, achieving the highest average among open models and best results on three of five spatial reasoning benchmarks.

By Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An, Haodong Jing, Ruqi Huang