arXiv:2604. 25834v2 Announce Type: replace Abstract: With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content consumption platforms.
By Wenhao Li, Zihan Lin, Zhengxiao Guo, Jie Zhou, Shukai Liu, Yongqi Liu, Chuan Luo, Chaoyi Ma, Ruiming Tang, Han Li
arXiv:2603. 22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.
By Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu
The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv:2606. 19627v1 Announce Type: cross Abstract: The digital commerce landscape is shifting from static, search-driven catalogs to dynamic, immersive video feeds.
By Katya Mirylenka, Egor Malykh, Mahdyar Ravanbakhsh, Michael Gygli, Marco-Andrea Buchmann, Andrew Dzhoha, Svitlana Borzenko, Francesca Catino, Mohamed Gaafar, Maarten Versteegh, Thomas Kober, Dario d'Andrea, Ellie Langhans
The paper tackles the challenge of predicting student engagement from online tutoring videos, noting that engagement is a complex, multidimensional construct influenced by behavioral, emotional, and cognitive states. By analyzing the CASED dataset, the authors highlight the difficulty posed by high inter‑person variability and subjective annotations. They propose a multimodal framework that fuses implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral cues such as head pose, gaze, facial action units, emotion, and wavelet‑based audio features, integrating them via a Perceiver IO bottleneck and modeling participant personalities with variational posteriors. The system employs evidential regression and spectral‑normalized Gaussian process classification heads to provide uncertainty‑aware predictions, achieving competitive performance on the CASED challenge test set while offering well‑calibrated uncertainty metrics.
By Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
The paper introduces DUMoE, a drift‑aware multimodal user representation framework that models user preferences over time by integrating static profiles, short‑term signals, and long‑term dependencies. It employs a sparse mixture‑of‑experts interest adapter, where each expert captures a distinct latent interest and a gating network selects relevant experts for each user. A three‑stage training strategy decouples backbone learning, expert specialization, and gating optimization, and experiments on real social media data demonstrate that DUMoE outperforms existing methods in user interest and interaction prediction.
By Ziqing Qian, Haohang Chen, Shengqi Dang, Yuhan Xiong, Canyu Shen, Jiaying Lei, Nan Cao
ProtoFlow is a new multivariate time series forecasting framework that combines vector‑quantized autoencoding with prototype‑guided flow matching. It maps sequences into a discrete latent space, constructs a structured prior from the learned VQ codebook, and trains a DiT‑based rectified flow to transport samples from this prior to future latent representations conditioned on past observations. By replacing generic Gaussian noise with a learned prototype prior, ProtoFlow eliminates autoregressive rollout mismatch and achieves faster training convergence while delivering superior forecasting performance on benchmark datasets.
By Shibo Feng, Wanjin Feng, Yang Qiu, Deheng Ye, Peilin Zhao, Chunyan Miao
arXiv:2608. 16274v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as autoregressive item sequences.
By Pengfei Jia, Jingjian Wang, Jingmao Li, Ge Zhang, Feng Shi
arXiv:2608. 19735v1 Announce Type: new Abstract: We introduce RecPFN, a prior-fitted network that brings in-context learning to sequential recommendation.
By En Zhi Tan, Jia Xiang Lim, Bryan Lijie Chew, Tze Minh Ng, Benjamin Yan Han Yap
The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A...
Understanding user preferences from noisy and temporally evolving social media behaviors is fundamentally challenging due to interest drift, where user preferences shift across time and exhibit both m...
arXiv:2607. 29213v1 Announce Type: cross Abstract: Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent.
By Jiping Liu, Zhongmin Zhang, Zisen Sang, Zhijia Fang, Tao Ouyang, Ma Jiang, Shaopeng Liang, Zeyang Hou, Guodong Cao, Jia Jia