Diffusion and generative media

Image, video and audio generation — diffusion models, flow matching and the systems built on top of them.

2,813 stories · RSS feed

arXiv AI
Jul 31

WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models

arXiv:2607. 26621v2 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their adoption as backbones for foundation recommendation models (FRMs).

By Hao Jiang, Peiru Du, Pengfei Yao, Mengting Li, Siyuan Lou, Kuo Cai, Sheng Yu, Qiang Luo, Jian Liang, Ruiming Tang, Fei Pan, Peng Jiang, Wenwu Ou
arXiv Machine Learning
Jul 31

A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation

arXiv:2607. 27501v1 Announce Type: new Abstract: We present a lightweight approach to foundation modeling (\textbf{NEXUS}) that leverages pre-trained learning from collider physics data towards out-of-domain tasks in other scientific datasets, using a fully connected autoencoder model with approximately 3 million parameters.

By Liangyu Wu, Qibin Liu, Alexander Yue, Julia Gonski
arXiv Machine Learning
Jul 31

ARES: Anomaly Recognition Model For Edge Streams

arXiv:2511. 22078v2 Announce Type: replace Abstract: Many real-world scenarios involving streaming information can be represented as temporal graphs, where data flows through dynamic changes in edges over time.

By Simone Mungari, Albert Bifet, Giuseppe Manco, Bernhard Pfahringer
arXiv Machine Learning
Jul 31

Dynamically Scaled Activation Steering

arXiv:2512. 03661v2 Announce Type: replace Abstract: Activation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation.

By Alex Ferrando, Xavier Suau, Jordi Gonz\`alez, Pau Rodriguez
Hugging Face Trending Papers
Jul 30

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs' retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR).