Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Computer Vision
6d ago

Form and Void: Entangled Composition through an Autonomous AI Agent

The paper introduces FaV-A, a multimodal agent that generates positive and negative space compositions in a staged manner. It first creates a base object, analyzes its shape to identify negative‑space semantics, and then produces compositional instructions for the final image. Experiments show that this approach yields more visually coherent and semantically aligned compositions than direct zero‑shot multimodal large language model baselines.

By Shiwen Wang, Jian Yang, Xu Wang, Xincan Wang, Weiming Dong
arXiv AI
6d ago

Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

The paper proposes a new architecture for large language model (LLM) agents that enhances spatial understanding by combining geometrical tools with an LLM orchestrator in grid‑world environments. It first gathers geodesic trajectories, vector‑quantizes them to create a representative subset, and then has the LLM label each trajectory with a natural language description, turning them into reusable tools. During operation, the LLM selects the appropriate tool based on the current state and goal, while low‑level control executes the chosen trajectory, enabling efficient decision‑making in a partially observable 2D grid setting.

By Gabriel Turinici
arXiv AI
6d ago

MemFit: Efficient Long-Term Agentic Memory

MemFit is a long‑term memory system designed for conversational agents that stores each dialogue turn verbatim in an append‑only store, enabling near‑instantaneous, LLM‑free insertion. It indexes turns using segment summaries and employs an LLM‑free, multi‑path retrieval strategy that blends lexical and semantic signals with cross‑encoder reranking over caption‑augmented episodes. Experiments on LoCoMo, MemGallery, and LongMemEval‑S demonstrate state‑of‑the‑art performance while drastically reducing memory construction time and cost.

By Mitchell Piehl, Muchao Ye
arXiv AI
6d ago

Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation

The paper introduces Frame‑Evidence Co‑Adaptation (FECA), an evidence‑adaptive visual planning framework for generating analytical charts in multimodal deep research. FECA treats chart generation as an iterative interaction between visual frames and retrieved evidence, allowing frames to be guided, revised, or dropped based on evidence availability. Experiments on 100 real‑world research topics demonstrate that FECA improves numerical fidelity while maintaining report quality and chart utility.

By Yuxin Yue, Yingchen Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Xueqi Cheng
arXiv AI
6d ago

Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts

The paper investigates whether multimodal large language model (LLM) agents form partner‑specific conceptual pacts during repeated reference games. By comparing real agent dyads with a novel pseudo‑dyad baseline that preserves task structure but removes shared partner history, the authors show that humans compress descriptions and align labels through entrainment, whereas agents produce verbose, fixed‑effort descriptions with near‑ceiling label overlap regardless of partner. Thus, LLMs succeed in coordination without forming compact, history‑dependent referring expressions typical of human dialogue.

By Po-Ya Angela Wang, Chinmaya Mishra, Asl{\i} \"Ozy\"urek, Paula Rubio-Fern\'andez, Esam Ghaleb
arXiv AI
6d ago

When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies

The paper introduces TRUST, a token‑level reward model that predicts the correctness of partial chain‑of‑thought (CoT) traces in vision‑language‑action (VLA) policies, enabling monitoring and selective steering of reasoning. On driving and manipulation VLA benchmarks, TRUST improves reasoning accuracy and reduces collision rates and trajectory errors, though its impact on overall task performance varies across tasks. The study defines two evaluation axes—correctability and actionability—to assess when CoT can serve as a runtime safety interface.

By Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal
arXiv AI
6d ago

Personalized Image Generation with Reasoning and Reflection

The paper introduces a unified benchmark for personalized image generation that uses users' historical data—such as reviews, posts, images, captions, and metadata—to create images aligned with their lifestyle and aesthetic preferences. It defines two tasks: Personalized Scene Generation, which places objects in scenes reflecting user preferences for product presentation, and Personalized Creative Generation, which produces novel images faithful to a user's aesthetic for social media content. The authors also propose PEARL, a method that interleaves multimodal reasoning with a frozen image generator, achieving a 15% average improvement over baselines on personalization metrics.

By Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu, Yu Wang, Ryan A. Rossi, Tyler Derr
arXiv AI
6d ago

Geometric Similarity in VLM Low-Level Vision Representations

The paper introduces GeoSim, a four‑level framework for analyzing how vision‑language models (VLMs) represent low‑level vision tasks. It evaluates hidden‑layer representations across 24 tasks and two VLM paradigms—autoregressive models and diffusion transformers—using global similarity, local geometry, sparse feature decomposition, and topological verification. The study uncovers the organizing principles of low‑level visual representations and highlights their limitations in cross‑task and cross‑model agreement, offering an interpretability lens for assessing latent transferability and diagnosing model‑specific issues.

By Shao-Jun Xia, Huixin Zhang, Zhen Lei, Anlan Sun, Yuner Zhang, Xiaoyang Chen
arXiv AI
6d ago

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

The paper introduces Divide-and-Remember (D&R), a recursive memory method for vision-language-action (VLA) policies that optimises memory by maximizing the conditional mutual information between actions and memory given observations. D&R recursively divides the full history into top‑K selections over 2K tokens, using a shared lightweight selector across all recursion blocks to handle unbounded histories efficiently. Evaluated on the RoboMME benchmark of 16 long‑horizon manipulation tasks, D&R achieves state‑of‑the‑art success rates with consistent gains across all suites while using only 64 tokens, and similar improvements are observed in real‑robot experiments.

By Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh
arXiv AI
6d ago

CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

CineMR is a vision‑language model that integrates cardiac image‑analysis tools to perform quantitative assessment of cine cardiac MRI. It uses supervised fine‑tuning and Group Relative Policy Optimization to learn reliable tool invocation, achieving significantly higher accuracy on a multi‑cohort benchmark than existing medical VLMs. The model demonstrates that tool‑augmented reasoning improves ventricular measurement accuracy by up to 23.7% and is essential for robust quantitative CMR interpretation.

By Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang
arXiv AI
6d ago

ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning

ReCast is a plug‑in framework that protects private inputs for fixed‑interface multimodal reasoning by locally converting them into a shared textual evidence‑query record, rewriting entities and topics with a distilled model, and mapping numerical values through an invertible, role‑aware map. A reconstruction agent then generates the required media from this protected record, allowing a remote solver to return a program whose operands are restored locally before execution. On 4,000 held‑out ChartQA and NMSQA examples, ReCast achieves 75.10% accuracy, retaining 92.43% of the unprotected remote accuracy, and flags source‑content leakage in 7.95% of solver‑bound requests, outperforming all evaluated local baselines.

By Bingchen Pei, Lichong Chen, Bingxi Zhao, Ziang Wu, Sirui Wang, Min Zhang, Yanhao Chen, Qingxu Liu, Qiang Gao, Chang-Tien Lu, Bo Gao
arXiv AI
6d ago

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

Cog-VADU is a training‑free framework that transforms video anomaly detection into a sequential cognitive reasoning task. It uses Chain‑of‑Anomaly Detection Thought Prompting (CoADTP) to create a recurrent reasoning chain across video segments, preserving temporal memory and distinguishing complex anomalies from high‑motion normal activities. A cross‑modal re‑ranking stage aligns textual rationales with visual embeddings to enforce semantic consistency and temporal coherence, yielding competitive zero‑shot performance on multiple VAD benchmarks.

By Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais
arXiv AI
6d ago

VETO: Video Efficient Token Optimization for Vision Language Models

VETO (Video Efficient Token Optimization for Vision Language Models) is a plug‑in that reduces the quadratic cost of visual tokens in long‑video inference by applying dual‑axis compression: an intra‑frame compressor merges semantically similar tokens within each frame, and an inter‑frame compressor merges temporally redundant frames. By first compressing spatial dimensions, VETO lowers the cost of subsequent global temporal matching, surpassing single‑axis methods and achieving up to 45% faster inference on models such as LLaVA‑OneVision‑7B while maintaining or improving accuracy. The approach is universally applicable across LLaVA‑OneVision, InternVL‑2.5, and LongVA, preserving or enhancing zero‑shot accuracy even under extreme token budgets.

By Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen, Winston H. Hsu
arXiv AI
6d ago

MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation

arXiv:2505.17613v2 Announce Type: replace Abstract: Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation...

By Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu
arXiv Machine Learning
6d ago

XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning

arXiv:2607.04171v4 Announce Type: replace-cross Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...

By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng