arXiv AI By Jialin Liu, Zhaorui Zhang, Ray C. C. Cheung

Beyond Modality Harmony: Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation

Read the original on arXiv AI →

The paper introduces OrthoRec, a multimodal recommender system that challenges the common "modality harmony" assumption by addressing conflicts between multimodal features and collaborative patterns. It employs Collaborative‑Guided Orthogonal Purification (CGOP) to separate useful multimodal signals from noisy orthogonal components, and a Topology‑Aware Routing Mixture‑of‑Experts (TAR‑MoE) to adaptively integrate purified modalities based on collaborative topology. Experiments on Amazon datasets demonstrate that OrthoRec outperforms recent baselines and remains robust to modality noise and item sparsity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 7

MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning

MURAL is a multimodal recommendation framework that replaces static similarity graphs with a dynamic topology discovery process. It uses an Adaptive Edge Learner to find latent item-item correlations efficiently and an Uncertainty-Aware Fusion module to down‑weight noisy modality signals based on aleatoric uncertainty. The model also incorporates a contrastive teacher‑student alignment to stabilize training and has been shown to outperform state‑of‑the‑art baselines on large‑scale TikTok and Amazon datasets, providing both higher accuracy and interpretability.

By Ahmad Mousavi (Department of Mathematics,Statistics American University), Majid Alikhani (Independent Researcher), Yeon-Chang Lee (Department of Computer Science,Engineering Ulsan National Institute of Science,Technology), Roberto Corizzo (Department of Computer Science American University), Yeganeh Abdollahinejad (Department of Biosystems,Agricultural Engineering Michigan State University)
arXiv AI
Aug 13

VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation

arXiv:2603. 17450v2 Announce Type: replace-cross Abstract: Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity and prevents Collaborative Filtering (CF) signals from being fully integrated into item representations.

By Junyoung Kim, Woojoo Kim, Wonbin Kweon, Jaehyung Lim, Dongha Kim, Hwanjo Yu
arXiv Machine Learning
Aug 17

MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity

arXiv:2608. 13911v1 Announce Type: new Abstract: Federated multimodal medical AI faces modality heterogeneity at both the client and sample levels: clients may systematically lack access to specific modality types, while individual records within the same client may contain different partial modality subsets.

By Adiba Orzikulova, Dong Min Kim, Jaehong Yoon, Sung-Ju Lee
arXiv Machine Learning
Jul 20

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.

By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang