arXiv AI

Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction

arXiv:2606. 27539v1 Announce Type: cross Abstract: Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations.

arXiv AI
Aug 19

M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction

M3TR is a temporal retrieval‑enhanced multi‑modal framework for predicting micro‑video popularity. It introduces a Mamba‑Hawkes Process module to model user feedback as self‑exciting events, capturing long‑range temporal dependencies. A temporal‑aware retrieval engine then identifies historically relevant videos by combining multi‑modal content similarity with popularity trajectory similarity, augmenting the target video’s features for improved prediction accuracy.

By Jiacheng Lu, Weijian Wang, Mingyuan Xiao, Yang Hua, Tao Song, Bo Peng, Cheng Hua, Haibing Guan
arXiv Machine Learning
Aug 27

Drift-Aware Multimodal User Representation Learning via Multi-Scale Temporal Modeling and Sparse Mixture-of-Experts

The paper introduces DUMoE, a drift‑aware multimodal user representation framework that models user preferences over time by integrating static profiles, short‑term signals, and long‑term dependencies. It employs a sparse mixture‑of‑experts interest adapter, where each expert captures a distinct latent interest and a gating network selects relevant experts for each user. A three‑stage training strategy decouples backbone learning, expert specialization, and gating optimization, and experiments on real social media data demonstrate that DUMoE outperforms existing methods in user interest and interaction prediction.

By Ziqing Qian, Haohang Chen, Shengqi Dang, Yuhan Xiong, Canyu Shen, Jiaying Lei, Nan Cao
arXiv Machine Learning
Jun 5

Bridging the Semantic-Collaborative Gap: An Asymmetric Graph Architecture for Cold-Start Item Recommendation

arXiv:2606. 06225v1 Announce Type: cross Abstract: Collaborative filtering and graph-based recommendation models are highly effective because they leverage observed user interactions, but this dependence creates a fundamental cold-start challenge when newly added content has no interaction history.

By Anh Truong, John Trenkle, Yuanbo Chen, Honghong Zhao, Abdullah Alchihabi, Effy Fang, Michael Tamir
arXiv Machine Learning
Jul 20

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.

By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv AI
Sep 10

LLMs for Social Network Modeling: From Network Generation to Dynamic Processes

The paper reviews how large language models (LLMs) are being used to model social networks, highlighting their ability to represent users, relationships, and interactions through natural language. It categorizes existing work into network generative models—split into selection‑based and interaction‑based approaches—and dynamic process models, which cover opinion dynamics, information diffusion, and rumor propagation. The survey also discusses the advantages of LLMs for realistic, context‑aware social behavior, while noting limitations such as social biases and prompt sensitivity, and outlines open research challenges and future directions.

By Shikha Mallick, Alex Thomo, Akrati Saxena