Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.
arXiv:2608. 10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot.
Doc‑CoB introduces a Chain‑of‑Boxes framework that enhances document understanding by progressively focusing on query‑relevant layout regions while preserving global context. It selects key layout boxes and then applies visual prompting for deeper analysis, supported by two new reasoning tasks and an automatic pipeline that generates 249k training samples with intermediate visual supervision. Experiments across seven benchmarks and four popular models demonstrate significant performance gains, underscoring the method’s effectiveness and broad applicability.
The paper titled "Visual General Intelligence: A White Paper" reexamines intelligence from a vision-centered perspective, questioning whether visual experience and learning can lead to artificial general intelligence (AGI). It compares the success of language models like GPT, which transfer to unseen tasks via autoregressive modeling on large text corpora, with the potential of visual modalities such as images, videos, and geometry to develop similar capabilities. The authors aim to outline principles for computer vision in the AGI era, including input modalities, benchmarks, learning paradigms, and the interplay between vision and other modalities like language.
arXiv:2601. 13591v2 Announce Type: replace Abstract: Recent LLM-based data agents aim to automate data science tasks ranging from data analysis to deep learning.