arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.
By Xinpeng Dong, Min Zhang, Kairong Han, Xu Tan, Fei Wu, Kun Kuang
arXiv:2608. 19208v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored.
By Yinfeng Wang, Zhiyuan Yao, Zheren Fu, Lei Zhang, Zhendong Mao
arXiv:2605.26380v2 Announce Type: replace-cross
Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. How...
By Jingru Chen, Yiming Liu, Mingtao Chen, Sijie Chen, Richeng Xuan, Liang Yang, Zhichao Hu, Fanyang Lu
ASCIIBench is a new benchmark that evaluates large language models on generating and classifying ASCII-text images, using a dataset of 5,315 labeled ASCII images. The authors also release a fine‑tuned CLIP model adapted to capture ASCII structure for evaluation. Their analysis shows that cosine similarity on CLIP embeddings fails to separate most categories, indicating a representation bottleneck rather than generational variance.
By Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu
arXiv:2606. 01022v1 Announce Type: cross Abstract: Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value for domains such as marketing, advertising, and E-commerce.
By Zhihong Liu, Siqi Kou, Zheng Li, Ye Ma, Quan Chen, Peng Jiang, Kai Yu, Zhijie Deng
arXiv:2509. 12159v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models have demonstrated exceptional performance in UI2Code tasks, significantly enhancing website development efficiency.
By Jingyu Xiao, Zhongyi Zhang, Yuxuan Wan, Yintong Huo, Yang Liu, Michael R. Lyu
The paper investigates how different multimodal design choices affect the performance of misinformation detection systems. Using over 3,375 experiments across three benchmark datasets and various pre‑trained vision and language models, the authors systematically compare design options and conduct robustness analyses. The study offers practical guidance on which choices improve detection, when they may fail silently, and which pipeline components most influence model behavior, addressing four key research questions.
By Akshit Sharma, Prashant W. Patil
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page layout and cross-page visual coherence.
arXiv:2608.28696v1 Announce Type: new
Abstract: Visual in-context learning (ICL) with multimodal large language models (MLLMs) is effective for fine-grained visual classification, but each retrieved...
By Hardik Jindal, Soumyabrata Pal, Sayak Ray Chowdhury
LensVLM is an inference framework and post‑training recipe that lets Vision‑Language Models (VLMs) process compressed images of text by selectively expanding only the relevant parts back to full resolution. Using Qwen3.5‑9B‑Base, LensVLM achieves accuracy comparable to full‑text models at 4.3× compression and outperforms other compression baselines up to 10.1× across seven text QA benchmarks, while also improving performance on multimodal document and code tasks as compression increases.
By Roy Xie, Dan Friedman, Donghan Yu, Bowen Pan, Christopher Fifty, Jang-Hyun Kim, Xianzhi Du, Zhe Gan, Vivek Rathod, Bhuwan Dhingra
The paper introduces AttWarp, a lightweight technique that uses a multimodal large language model’s cross‑modal attention to perform rectilinear warping of input images at test time. By reallocating spatial resolution toward query‑relevant regions without altering model weights or architecture, AttWarp preserves global context while making small objects and subtle relationships easier for the model to read. Experiments on five benchmarks and four MLLMs show consistent accuracy gains, improved compositional reasoning, and reduced hallucinations compared to baseline image‑manipulation methods.
By Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra, Jeonghwan Kim, Madhav Kanda, Hyeonjeong Ha, Svetlana Lazebnik, Heng Ji, Unnat Jain
The paper surveys Multimodal Code Intelligence, focusing on tasks where code is generated, edited, refined, or reasoned about under visually grounded inputs such as screenshots, charts, and videos. It categorizes the field by the role of code—rendered artifact, editable structure, intermediate reasoning trace, or executable tool interface—and organizes benchmarks into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. The authors argue that reliable evaluation must include evidence of semantics and interaction beyond visual fidelity, and propose four verification-centered research directions to advance the field toward evidence-grounded executable systems.
By Xuanle Zhao, Qiushi Sun, Jingyu Xiao, Xuexin Liu, Haoyue Yang, Qiaosheng Chen, Xianzhen Luo, Jing Huang, Yufeng Zhong, Lei Chen, Shuai Fu, Zhenlin Wei, Jinhe Bi, Lei Jiang, Haibo Qiu, Siqi Yang, Peng Shi, Jian Hu, Zhixiong Zeng