The paper introduces the Robust Graph Clustering Network for Multiple Missing Data (RGCN), a method designed to cluster graphs with simultaneous missing node attributes and structural links. RGCN employs a view‑decoupled dual‑branch imputation to reduce cross‑view interference, a multi‑hyperspherical mixture prior to improve cluster compactness and separability on a directional latent manifold, and a boundary‑aware contrastive enhancement objective to counteract cluster blurring caused by imputation bias. Experiments on real‑world datasets show that RGCN consistently outperforms state‑of‑the‑art baselines across various missing data patterns.
By Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni
The paper introduces the concept of protocol divergence, showing that identical nominal missing rates can lead to vastly different learning regimes in incomplete multi‑view clustering. It critiques existing evaluation practices that ignore observation structure and proposes CRAFT, a train‑once framework that fuses observed views with mask‑aware attention, enabling efficient deployment across multiple missing‑view protocols. Experiments on CUB, MultiFashion, and other benchmarks demonstrate CRAFT’s superior performance and significant computational savings through checkpoint reuse.
By Haolu Liu, Xiyue Wang, Xuanting Xie, Liangjian Wen, Zhao Kang
arXiv:2607. 27761v1 Announce Type: new Abstract: In recent years, multi-view clustering has attracted widespread research interest.
By Shubin Ma, Liang Zhao, Chuanye He, Zhenjiao Liu, Liang Zou, Lin Yuanbo Wu, Yu Shao
The paper presents a semi‑supervised generative model for multi‑view learning that handles missing views and missing labels. It combines a likelihood‑based approach for unlabeled data with an information bottleneck (IB) framework for labeled data, incorporating modality‑specific information and cross‑view mutual information maximization to learn a shared latent space. Experiments show improved predictive and generative performance on complex datasets with limited labeled samples.
By Yiyang Shen, Weiran Wang
The paper introduces SIMPLE, a prior‑fitted multi‑view in‑context learner that learns a reusable, task‑conditioned inference procedure instead of a fixed fusion function. By generating synthetic task priors in embedding space, SIMPLE can handle diverse view configurations, class structures, and missingness patterns. Experiments on multi‑view and multi‑omics benchmarks show that a frozen SIMPLE model performs competitively, and lightweight adapter calibration further improves performance across most datasets.
By Jielong Lu, Zhihao Wu, Jiajun Yu, Zhaoliang Chen, Haishuai Wang
arXiv:2609.15305v1 Announce Type: cross
Abstract: Urban region representation learning commonly combines heterogeneous data sources, such as mobility flows, points of interest, and land-use informati...
By Sean Bin Yang, Ying Sun, Zongyi Xu, Tung Kieu, Jilin Hu, Bin Yang, Kristian Torp, Hua Lu, Torben Bach Pedersen
arXiv:2504. 18455v2 Announce Type: replace-cross Abstract: We study distributed multiview representation learning, a problem in which $K$ clients each observe a distinct but possibly statistically correlated view.
By Milad Sefidgaran, Piotr Krasnowski, Abdellatif Zaidi
MVVBench is a new benchmark for multi‑view video reasoning that tests vision‑language models on tasks requiring integration of spatial and temporal evidence across multiple, often non‑overlapping camera streams. The benchmark contains questions that cannot be answered from any single view or single moment, forcing models to jointly reason across views and time. It evaluates six capabilities—including attribute identification, relative distance, camera pose, and compositional counting—and provides human‑authored QA, rigorous verification, and detailed error analysis.
"whyItMatters":"The benchmark offers a rigorous evaluation of 4D multi‑view reasoning and a foundation for future progress toward reliable embodied perception."
By Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi, Byung-Hoon Kim
arXiv:2606. 09301v1 Announce Type: new Abstract: Multimodal federated graph learning (MM-FGL) aims to collaboratively learn from decentralized graphs with text and images.
By Zekai Chen, Miao Zhang, Jiayang Xing, Xunkai Li, Xun Wu, Rong-Hua Li, Guoren Wang
arXiv:2610.01098v1 Announce Type: new
Abstract: Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D recons...
By Hanyuan Xiao, Gonglin Chen, Haolin Xiong, Wenbin Teng, Haiwei Chen, Yajie Zhao
The paper introduces MVLGeo, a unified framework for cross-view object geo-localization that combines multiple viewpoints into a single model. It employs Vision‑Language Reranking to use contextual text from the query view, a multi‑view Mixture‑of‑Experts architecture to share knowledge and reduce redundancy, and an adaptive elliptical prior for positional encoding. Experiments on CVOGL benchmarks show that MVLGeo achieves state‑of‑the‑art performance and robustness to input degradation.
By Xuyu Fan, Qi Ming, Zhu Han, Liuqian Wang, Siyuan Cao, Xiaohan Zhang, Xudong Zhao, Mingjing Zhao, Yuhan Zhang
arXiv:2606. 05109v1 Announce Type: new Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information.
By Vasiliki Rizou, Pascal Frossard, Dorina Thanou