Hugging Face Blog

Going multimodal: How Prezi is leveraging the Hub and the Expert Support Program to accelerate their ML roadmap

arXiv Computation and Language
Sep 23

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real‑world multimodal productivity, offering enhanced multimodal understanding, reasoning, and long‑horizon agentic task performance. It builds on a sparse mixture‑of‑experts architecture, extends its context window to one million tokens, and supports long‑context multimodal reasoning and planning. The release includes Qwen-MM-Plugins for native audio and video support and Qwen-Live-Harness for building responsive, real‑time multimodal agents, with extensive evaluations confirming strong performance across multimodal tasks.

By Qwen Team
arXiv AI
Sep 1

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

arXiv:2508.13186v2 Announce Type: replace-cross Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However...

By Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu, Jun Dong, Haoyang He, Hao Lu, Haozhe Zhang, Chenchen Jing, Zhen Li, Chuanhao Li, Jiayi Tian, Chenchen Zhang, Tianhao Peng, Yancheng He, Jihao Gu, Hui Huang, Donghao Zhou, Yuanxing Zhang, Jian Yang, Ge Zhang, Wenhao Huang, Zhaoxiang Zhang, Qiangpeng Yang, Shilei Wen
arXiv Computation and Language
Sep 28

Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward

The paper introduces UniSandbox, a decoupled evaluation framework with controlled synthetic datasets, to study whether understanding informs generation in Unified Multimodal Models. Results show a notable understanding‑generation gap, especially in reasoning generation and knowledge transfer. Explicit Chain‑of‑Thought (CoT) in the understanding module bridges this gap, and self‑training can internalize CoT for implicit reasoning during generation; query‑based architectures also exhibit latent CoT‑like properties that aid knowledge transfer.

By Yuwei Niu, Weiyang Jin, Jiaqi Liao, Chaoran Feng, Peng Jin, Bin Lin, Zongjian Li, Bin Zhu, Weihao Yu, Li Yuan
arXiv Computation and Language
Sep 18

Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion

The paper introduces TrioRAG, a graph-free multimodal retrieval-augmented generation framework that combines evidence from the question, an anchor image, and a VLM-enhanced query via late fusion. It also presents AutoQA, a benchmark featuring noisy web-sourced images that require reasoning across manuals. TrioRAG outperforms graph-based systems on three benchmarks while cutting costs and speeding up inference by 1.6–2.3×.

By Tithi Rakshit, Hongkuan Zhou, Lavdim Halilaj, Yuqicheng Zhu
arXiv AI
Aug 24

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

The paper surveys how large multimodal models (LMMs) enhance agentic frameworks that combine perception, memory, reasoning, planning, and action. It examines the integration of multiple modalities—text, images, audio, and video—through delegated, late‑fusion, and early‑fusion architectures, and maps these designs to agent capabilities. The survey also reviews multimodal agentic systems in robotics, web navigation, multimedia content creation, and video understanding, evaluating performance, efficiency, and scalability trade‑offs.

By Neel Mokaria, Rishie Raj, Dheeraj Baiju, Xiaoqian Shen, Shraman Pramanick, Kevin Qinghong Lin, Arda Senocak, Mike Zheng Shou, Philip Torr, Mohamed Elhoseiny, Yapeng Tian, Ruohan Gao, Salman Khan, Sayan Nag, Sanjoy Chowdhury, Dinesh Manocha
arXiv AI
4d ago

MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation

arXiv:2505.17613v2 Announce Type: replace Abstract: Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation...

By Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu
arXiv AI
Sep 18

From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning

The survey "From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning" reviews over 300 works on efficient multimodal learning (EML), proposing a structured taxonomy that spans model, algorithm, and system layers. It synthesizes how cross‑layer co‑design addresses the Efficiency‑Utility‑Privacy trade‑off and illustrates this through a case study of multimodal large language models. The paper also offers optimization blueprints for various domains, discusses a shift toward self‑regulating intelligence, and outlines open challenges for future EML research.

By Pan Wang, Siwei Song, Hui Ji, Siqi Cao, Heng Yu, Zhijian Liu, Huanrui Yang, Yingyan Celine Lin, Beidi Chen, Mohit Bansal, Xiaoming Liu, Pengfei Zhou, Ming-Hsuan Yang, Tianlong Chen, Jingtong Hu
arXiv AI
Aug 11

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

arXiv:2608. 09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion.

By Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren