The paper investigates how affective fine‑tuning shapes the internal architecture of multimodal foundation models. By analyzing 13 model instances across nine designs, it finds that adapting the feed‑forward network (FFN) consistently outperforms attention‑only adaptation and nearly matches full‑model tuning, revealing the FFN as an efficient adaptation substrate. Moreover, joint optimization leads to emergent functional specialization, notably a prominent gate projection pathway, which the authors exploit in Gate‑Focused Efficient Tuning (GET) to achieve 96.2–98.0% of full‑model performance with only 19.3–24.5% of the trainable parameters.
By Zhen Zhang, Runhao Zeng, Sicheng Zhao, Xiping Hu
This thesis introduces methods for assessing and enhancing the robustness of large language models (LLMs) against adversarial input variations. It proposes a generative robustness metric, R_stab(f), based on Jensen‑Shannon divergence, and proves bounds for localized attacks. The work also presents adaptive attacks, Trojan detection techniques, defense strategies for multi‑layer systems, and lightweight attestation for Model Context Protocol (MCP) agentic systems, all implemented in the JudgeGuard, TrojanArmor, and MCPSec suites.
By Narek Maloyan
The study investigates how two computational dimensions—model depth and refinement steps—affect intelligibility and speaker identity in masked-diffusion text‑to‑speech systems. Experiments with 15 models (19–133 M parameters) and up to 16 refinement steps show that refinement improves intelligibility more than identity, with a 1.86× asymmetry that persists even after retraining. Best‑of‑K search can recover identity when refinement fails, and analysis indicates that depth and steps target distinct bottlenecks, requiring separate optimization.
By Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh
The paper introduces ImmuneAgent, a closed‑loop AI system that combines multimodal reasoning, continual meta‑learning, and wet‑lab feedback to identify broadly neutralizing antibodies (bnAbs) from human B cell repertoires. Applied to vaccinated or infected cohorts, ImmuneAgent achieved a 55% neutralization discovery rate and an 11% bnAb yield, outperforming existing sequence‑based predictors and co‑folding models. Five discovered antibodies provided full in vivo protection against lethal influenza, and the system uncovered conserved bnAb reservoirs and structural signatures that enabled cross‑viral antibody discovery without antigen‑specific sorting.
By Hantao Lou, Jianqing Zheng, Can Yue, Meihan Zhang, Yuanchao Bao, Yu Chen, Mengting Huang, Yupeng Yang, Qianyu Pan, Nana Fu, Yansong Shi, Hongli Li, Yangyang Chai, Ruyi Chen, Wansheng Li, Zhu Liang, Rongmei Yao, Yuanhan Mo, Lei Wang, Chunmei Wang, Yun Quan, Qiong Zhang, Xiangxi Wang, Xuetao Cao
arXiv:2610.02214v1 Announce Type: cross
Abstract: Ship structures govern vessel strength, safety, and manufacturability, but their design must satisfy hundreds of classification society requirements,...
By Noah J. Bagazinski, Md. Ferdous Alam, Jaya Manideep Rebbagondla, Faez Ahmed
arXiv:2610.02768v1 Announce Type: new
Abstract: Multimodal graph learning has recently emerged as an effective paradigm for in corporating inter-entity relationships into multimodal representations....
By Zekai Chen, Kai Hu, YuXin Zeng, Xunkai Li, Xun Wu, Yinlin Zhu, Zhengyu Wu, Xu Wang, Rong-Hua Li
arXiv:2610.03123v1 Announce Type: cross
Abstract: Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed...
By Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari, Samip Ghimire, Saroj Poudel, Binod Bhattarai, Danda Pani Paudel
arXiv:2610.03153v1 Announce Type: cross
Abstract: Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resource...
By Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen
arXiv:2610.02638v1 Announce Type: new
Abstract: Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which f...
By Jie Jin, Ziyin Ma, Min Yin, Jinyu Chen, Haigang Song, Zhikun Pang, Xiaowen Zhang
arXiv:2610.02824v1 Announce Type: new
Abstract: Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requireme...
By Yuxuan Fan, Jaehong Yoon
arXiv:2603.13856v3 Announce Type: replace
Abstract: Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the...
By Naaisha Agarwal, Yihan Wu, Xin Guan, Ayaan Garg, Yikuan Hu, Mohan Li, Vincenzo Collura, Wang-Zhou Dai, Yao-Xiang Ding, Emanuele Sansone
arXiv:2610.03036v1 Announce Type: cross
Abstract: We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge...
By Jiangang Han
arXiv:2610.03084v1 Announce Type: cross
Abstract: Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complem...
By Omar Elfatairy, Maria A. Bravo, Jessica Bader, Zeynep Akata
arXiv:2610.02507v1 Announce Type: new
Abstract: We present MeshQuery, a training-free agentic approach to automatic UV unwrapping of production-grade quad meshes. A Vision-Language Model (VLM) plans...
By Marco Schouten, Arthur Roullier, Elie Michel, Ruben Wiersma, Axel Paris, Tamy Boubekeur
arXiv:2610.02597v1 Announce Type: new
Abstract: Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowle...
By Zhenghao Zhao, Chi Zhang, Qingshuang Chen, Yelin Kim
arXiv:2610.03476v1 Announce Type: cross
Abstract: Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and...
By Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi
arXiv:2604.04692v3 Announce Type: replace-cross
Abstract: Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-o...
By Jaeyoon Jung, Yejun Yoon, Kunwoo Park
arXiv:2610.03002v1 Announce Type: new
Abstract: Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvem...
By Huijuan Wang, Chufan Shi, Cheng Yang, Yaokang Wu, Taylor Berg-Kirkpatrick, Xuezhe Ma
arXiv:2608.05369v2 Announce Type: replace-cross
Abstract: Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct r...
By Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu
arXiv:2610.02819v1 Announce Type: new
Abstract: Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics...
By Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Qimin Wu, Jingru Fan, Chen Qian, Yanfeng Wang, Yu Wang