Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Computer Vision
3d ago

Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization

The paper investigates how affective fine‑tuning shapes the internal architecture of multimodal foundation models. By analyzing 13 model instances across nine designs, it finds that adapting the feed‑forward network (FFN) consistently outperforms attention‑only adaptation and nearly matches full‑model tuning, revealing the FFN as an efficient adaptation substrate. Moreover, joint optimization leads to emergent functional specialization, notably a prominent gate projection pathway, which the authors exploit in Gate‑Focused Efficient Tuning (GET) to achieve 96.2–98.0% of full‑model performance with only 19.3–24.5% of the trainable parameters.

By Zhen Zhang, Runhao Zeng, Sicheng Zhao, Xiping Hu
arXiv Machine Learning
3d ago

Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

This thesis introduces methods for assessing and enhancing the robustness of large language models (LLMs) against adversarial input variations. It proposes a generative robustness metric, R_stab(f), based on Jensen‑Shannon divergence, and proves bounds for localized attacks. The work also presents adaptive attacks, Trojan detection techniques, defense strategies for multi‑layer systems, and lightweight attestation for Model Context Protocol (MCP) agentic systems, all implemented in the JudgeGuard, TrojanArmor, and MCPSec suites.

By Narek Maloyan
arXiv AI
3d ago

Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS

The study investigates how two computational dimensions—model depth and refinement steps—affect intelligibility and speaker identity in masked-diffusion text‑to‑speech systems. Experiments with 15 models (19–133 M parameters) and up to 16 refinement steps show that refinement improves intelligibility more than identity, with a 1.86× asymmetry that persists even after retraining. Best‑of‑K search can recover identity when refinement fails, and analysis indicates that depth and steps target distinct bottlenecks, requiring separate optimization.

By Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh
arXiv AI
3d ago

Multimodal reasoning for broadly neutralizing antibody discovery from label-free human B cell repertoires across virus families

The paper introduces ImmuneAgent, a closed‑loop AI system that combines multimodal reasoning, continual meta‑learning, and wet‑lab feedback to identify broadly neutralizing antibodies (bnAbs) from human B cell repertoires. Applied to vaccinated or infected cohorts, ImmuneAgent achieved a 55% neutralization discovery rate and an 11% bnAb yield, outperforming existing sequence‑based predictors and co‑folding models. Five discovered antibodies provided full in vivo protection against lethal influenza, and the system uncovered conserved bnAb reservoirs and structural signatures that enabled cross‑viral antibody discovery without antigen‑specific sorting.

By Hantao Lou, Jianqing Zheng, Can Yue, Meihan Zhang, Yuanchao Bao, Yu Chen, Mengting Huang, Yupeng Yang, Qianyu Pan, Nana Fu, Yansong Shi, Hongli Li, Yangyang Chai, Ruyi Chen, Wansheng Li, Zhu Liang, Rongmei Yao, Yuanhan Mo, Lei Wang, Chunmei Wang, Yun Quan, Qiong Zhang, Xiangxi Wang, Xuetao Cao
arXiv Machine Learning
3d ago

Can AI Understand the Language of Origami?

arXiv:2603.13856v3 Announce Type: replace Abstract: Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the...

By Naaisha Agarwal, Yihan Wu, Xin Guan, Ayaan Garg, Yikuan Hu, Mohan Li, Vincenzo Collura, Wang-Zhou Dai, Yao-Xiang Ding, Emanuele Sansone
arXiv Computer Vision
3d ago

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

arXiv:2608.05369v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct r...

By Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu