Hugging Face Blog

Launching the Artificial Analysis Text to Image Leaderboard & Arena

arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
arXiv AI
Sep 28

SlideLab: Audience-Centered Scientific Slide Generation and Evaluation

SlideLab is a training‑free, multi‑agent framework that generates scientific presentations directly from research papers. It first plans a coherent narrative, then iteratively builds and refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab outperformed both open‑source and commercial systems on 77% of papers while using about four times fewer inference tokens than the strongest open‑source baseline. The authors also introduce ConfArena, an audience‑oriented evaluation framework that simulates a conference room and assesses presentations slide by slide, matching human system rankings and detecting issues such as falsified numbers, degraded figures, dropped slides, and shuffled slide order.

By Vidushee Vats, Karun Sharma, Yuxia Wang
arXiv AI
Jun 2

Agentic-J: An AI Agent for Biological Microscopy Image Analysis

arXiv:2606. 02080v1 Announce Type: cross Abstract: Biological image analysis increasingly demands integration across heterogeneous tools, programming environments, and domain knowledge that few researchers can command simultaneously.

By Lukas Johanns, Marilin Moor, Davide Panzeri, Yu Zhou, Xinyi Chen, Nora F. K. Pauly, Zixuan Pan, Matthias Gunzer, Andreas M\"uller, Yiyu Shi, Hedi Peterson, Jianxu Chen
arXiv Machine Learning
Jul 27

Multi-Agent Debate and Visual Information Extraction for SeePhys Pro: A 1st-Place Technical Report from ICML 2026 AI4Math Track 3 Challenge

arXiv:2607. 21946v1 Announce Type: new Abstract: This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose statement and figure may be given partly or entirely as an image.

By Jiseok Kwak, Suhyeon Jo, Taewoo Kim, Yeongmin Kim, Byeonghu Na, Il-chul Moon