Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,813 stories · RSS feed

arXiv AI
Sep 30

FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design

arXiv:2609.36064v1 Announce Type: cross Abstract: Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly...

By Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida, Gregg Streuber, Farnoosh Javadi, Jeffrey Landes, Angela Ju, Muhammad Azam, Bryan Lim, Johan Luttun, Indrajeet Haldar, Jonathan Shaw, Beatriz Guerra, Ivan Sosnovik, James Stoddart, Robert Giaquinto, Adam Gaier
arXiv Computer Vision
Sep 30

Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving

Vision‑Language Models (VLMs) used in autonomous driving are evaluated on their ability to understand velocity through three tasks: surrounding‑agent speed, current ego speed, and short‑horizon future ego‑speed. Experiments on nuScenes show that while current ego speed can be accessed internally, verbal outputs are poor and surrounding‑agent speed is weakly encoded; temporal cues and frame order are largely ignored. Specialized driving models (Alpamayo‑1.5) improve latent speed representations but still leave gaps between internal knowledge and verbal readout, indicating that good planning outputs do not guarantee reliable dynamic state recovery.

By Katharina Winter, Stefan Englmeier, Fabian B. Flohr
arXiv Computer Vision
Sep 30

Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

The paper introduces Q-TOFC, a query‑guided task‑oriented visual feature compression method that uses residual vector quantization to encode merged features as compact codebook index sequences. By incorporating query relevance into feature aggregation and adding a quantization error compensation adapter, Q‑TOFC reduces visual payload by 53.6% compared to previous TOFC while preserving task performance. Experiments across seven multimodal benchmarks and latency tests confirm its effectiveness under bandwidth‑constrained uplinks.

By Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang, Yuan Shen
arXiv Computer Vision
Sep 30

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

EviViT is a lightweight attachment for pretrained vision transformers that learns where to focus detail in high‑resolution images. It uses human visual‑search traces to supervise a question‑conditioned evidence density, guiding regional re‑reading and efficient visual token allocation. The method connects regional features to the global scene via a sparse, coordinate‑aware bridge, improving fine‑grained accuracy across nine host models while using fewer tokens than global‑only processing.

By Yaoxin Niu, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, Ruqi Huang
arXiv AI
Sep 30

Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics

The paper benchmarks vision‑language models (VLMs) on two key tasks in connectomics: synapse detection (presence and polarity) and proofreading (split and merge errors). It evaluates 21 models (19 open, 2 closed) under zero‑shot, few‑shot, and LoRA fine‑tuning, comparing them to specialist models on datasets built from public resources. While most VLMs perform at chance in zero‑shot settings, LoRA fine‑tuning with a few thousand labels brings open models to specialist performance, and adapted VLMs outperform specialists on unseen species for merge‑error detection.

By Yicong Li, Junjie Wang, Leander Lauenburg, Ella Hugie, Alexandra Irger, Wanhua Li, Donglai Wei, Hanspeter Pfister
arXiv AI
Sep 30

MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms

MemEvo presents an automated approach to discover memory mechanisms for streaming video understanding. It frames memory design as a search over executable programs expressed in a lightweight domain-specific language, and uses a pretrained large language model to generate and refine candidate memory programs. The framework evaluates each candidate deterministically while keeping the underlying vision-language model frozen, ultimately producing a training-free, bounded-memory mechanism that shows strong performance on StreamingBench and OVO-Bench.

By Guohong Liu, Jialei Ye, Shanhui Zhao, Yunxin Liu, Yuanchun Li
arXiv AI
Sep 30

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

MatToolBench is a new benchmark that evaluates multimodal GUI agents on professional materials science software. It contains 204 tasks across 10 tools in three modalities—GUI operation, OriginPro scripting, and code-based database queries—executed inside a Windows 11 VM. The benchmark offers fine-grained, expert-decomposed scoring and a high-performing multimodal judge for aesthetic assessment, revealing that strong general benchmark performance does not transfer to scientific workflows.

By Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen
arXiv AI
Sep 30

Absorbed in Inertia: Activation Analysis for Computer-Use Agents

The paper investigates a phenomenon called inertia in computer‑use agents, where agents repeat ineffective actions despite recognizing their futility. By analyzing high‑dimensional activation states, the authors find that inertia corresponds to an absorbing region in activation space where values become stale. They propose a method called R$^3$—Reset, Reroute, Restore—to temporarily reset the agent’s context trajectory, escape the absorbing region, and then restore historical context, achieving a 17‑55% reduction in measured inertia across models.

By Giulio Segalini, Zhi Wen Soi, J\'er\'emie Decouchant, Lydia Chen
arXiv AI
Sep 30

VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict

VISTA (Value-Informed Semantic Trust Arbitration) is a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. It uses a log-odds decomposition to separate emotion expectation from cue diagnosticity, allowing appraisal to change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone, VISTA achieves 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency, and a frozen-backbone probe reaches 0.600 macro CCC for appraisal readout versus 0.505 for emotion-only fine-tuning.

By Jiale Dai, Liuxian Ma, Xiaoke Niu, Wenjing Zhang, Huiying Zhao, Zhaoxiang Liu, Shiguo Lian, Guojie Song
arXiv AI
Sep 30

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

SeekVLN is a new framework for Vision‑Language Navigation that addresses the problem of agents acting on insufficient evidence, termed Progress Myopia. It combines semantic progress reasoning with active evidence seeking, trained first with Future‑guided Reverse Generation to augment expert trajectories, and then refined via Counterfactual Contrastive Policy Optimization to reward beneficial seeking actions. Experiments on simulated benchmarks show significant gains, improving success rates by 12.7% on R2R‑CE and 7.5% on RxR‑CE, and real‑world tests demonstrate human‑like evidence‑seeking behavior.

By Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, Jingyan Jiang, Yaowei Wang, Zhi Wang
arXiv AI
Sep 30

Calibration-First Cross-Cohort Multimodal Temporal Learning for Transferable Asthma-Risk Forecasting

The paper introduces CALIBRA, a calibration-first multimodal temporal learning framework designed for transferable asthma‑risk forecasting across varying patient cohorts and sensor ecosystems. It processes multiple data streams—environmental, pulmonary, symptom, medication, wearable, and context—using dedicated recurrent encoders, a reliability‑conditioned gate, and gradient‑reversal training to mitigate cohort bias. CALIBRA employs a shrinkage‑based hierarchical logistic layer for probability calibration and split conformal prediction for abstention‑capable prediction sets, achieving competitive performance on a semi‑synthetic three‑cohort benchmark with controlled distribution shift.

By Taimoor Ahmad
arXiv AI
Sep 30

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.

By Aditya Sharma, Divya Saxena
arXiv AI
Sep 30

Render Before Reading: Visual Rendering as a Prompt Injection Defense

The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.

By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r
arXiv AI
Sep 30

Paired Multimodal Scaling Laws

The paper introduces a new multimodal scaling law that accounts for how the proportion of paired data influences loss in multimodal classification tasks. It shows that varying the pairing ratio under fixed total data budgets dramatically affects loss curves, with synergy only reducing when a critical threshold of paired data is reached. The proposed law decomposes total loss into four power-law components—redundancy, unique modality channels, and synergy—providing a more accurate fit (3.2% error) than existing models.

By Marcus Ma, Shrikanth Narayanan
arXiv AI
Sep 30

Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models

The paper introduces MATE, a reinforcement‑learning‑based post‑training framework for unified multimodal models that lets the generation and understanding branches challenge each other instead of cooperating. In MATE, each branch proposes candidate outputs that the other must reproduce, and the solver is trained on the worst‑handled candidate, creating an evolving adversarial loop without a separate adversary. Experiments on Janus‑Pro‑1B show that MATE improves generation and understanding metrics, including GenEval (+2.4), DPG‑Bench (+1.7), and an average of nine understanding benchmarks (+0.7), while enhancing consistency across image‑text cycles.

By Wentao Zhou, Weijie Gan, Jiayun Wang
arXiv AI
Sep 30

AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models

AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.

By Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi
arXiv AI
Sep 30

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.

By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li