A$^2$Safe is a framework for safe and effective Visual Question Answering that aligns counterfactual evidence with adaptive agent collaboration. It uses a Grounded Safety Evidence Board to make safety decisions explicit, enforcing invariance to safety‑irrelevant changes while allowing appropriate transitions when risk‑critical evidence changes. The system achieves a 95.72 SIUO safety score, reduces benign refusals on MOSSBench to 14.67%, and maintains a 78.34 average VQA score with 27.8% token overhead.
By Quanxing Xu, Ling Zhou, Xian Zhong, Jinyu Tian, Xiaohua Huang, Rubing Huang, Chia-Wen Lin
The paper introduces Bidirectional Reciprocal Learning (BRL), a parameter‑efficient fine‑tuning framework for referring image segmentation that operates on frozen vision foundation models. BRL employs two lightweight adapters—Reciprocal Attention Adapter (RAA) for token‑level cross‑modal attention and Reciprocal Gate Adapter (RGA) for channel‑level gating—to enable hierarchical, bidirectional information flow between vision and language. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show that BRL outperforms existing methods while updating fewer than 0.5% of backbone parameters.
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
VideoGen-Agent is a multimodal agent that uses multitask agentic reinforcement learning to coordinate external tools for video generation. It learns to augment, generate, and verify videos through multi‑turn interactions, guided by prompts and intermediate observations. On the new VABench benchmark, the agent improves base text‑to‑video performance by 19.1 points, and further upgrades to generation tools raise the score to 86.1, with human raters favoring the upgraded configuration in 84.3% of comparisons.
By Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin, Kaituo Feng, Suozhi Huang, Xiangyi Li, Yu Li, Chunyuan Li, Shilong Liu, Mengdi Wang
The paper introduces ProAction, a multimodal dataset of 10,000 samples comprising visual, audio, and text inputs across 12 daily-life scenarios, designed to support the Proactive Robot Action Reasoning (ProRobo) problem. It presents a two-stage human-in-the-loop annotation pipeline that incorporates appraisal and Theory-of-Mind considerations to generate cognitively grounded high-level action labels. The authors benchmark multimodal large language models and propose MMC2Act, showing that training on ProAction significantly improves proactive action reasoning compared to general-purpose models.
By Zhihao Gu, Kechao Zhu, Yuanfeng Wu, Mohan Liu, Ankit Kumar Shaw, ChenDong Hong, Xuanyu Chen, Dengchen Mei, Xu Tianyi, Lin Wang
The paper introduces GradCIR, a method for training composed image retrieval (CIR) systems on graded relevance rather than binary relevance. It uses a vision‑language model to generate queries and 4‑level relevance labels, an iterative feedback loop to mine hard negatives, and a hierarchy‑aware angular objective to directly optimize graded labels. Experiments on a Walmart catalog and FashionIQ show significant NDCG improvements and the system is deployed in Walmart’s live visual‑search traffic.
By Anubhav Gupta, Hrushikesh Mohapatra, Prijith Chandra, Asish Mohapatra, Anuj Garg, Arvind Maan, Sudip Datta, Venkat Bulusu, Sitesh Kumar Jalan
The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
The paper introduces CSC, a calibrated-simplicity framework for detecting social bots in the era of large language models. CSC combines a simplified prototype-guided graph expert, calibrated simplex-constrained fusion, and a lightweight inconsistency expert to address modality conflict between semantic and structural signals. Experiments on TwiBot-22, TwiBot-20, and MGStBot-large demonstrate that CSC improves calibrated decision quality while maintaining competitive performance across benchmarks.
By Yipeng Qian, Pengjie Zhao, Chaoxi Niu
arXiv:2609.22092v1 Announce Type: cross
Abstract: Electroencephalography (EEG) is a low-cost and non-invasive signal source for dementia screening, yet existing EEG-based studies remain difficult to...
By Haitian Wang, Chamara Madarasingha, Redowan Mahmud, Aneesh Krishna, Ryu Takechi
arXiv:2609.22684v1 Announce Type: cross
Abstract: Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA...
By Wenzhuo Li, Qiongfeng Shi, Yi Zhou
arXiv:2609.23038v1 Announce Type: cross
Abstract: Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requi...
By Kaixiang Yao, Xu Wang, Miao Pan, Hu Xiyue, Weishi Wang, Daniel Dahlmeier, Jintao Chen, Yongliang Shen, Xuhong Zhang, Wenqi Zhang
arXiv:2609.23376v1 Announce Type: cross
Abstract: Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sourc...
By Yi Xu, Cheng Chen, Wenzhuo Lei
arXiv:2609.23703v1 Announce Type: cross
Abstract: Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrat...
By Kemal Kirtac
arXiv:2609.14360v2 Announce Type: replace
Abstract: Reliable small-molecule identification often requires complementary evidence from multiple spectroscopic measurements. In practice, however, spectr...
By Bowen Gao, Lei Zhu, Yiying Wang, Wenjie Yu
arXiv:2411.05005v2 Announce Type: replace-cross
Abstract: Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, m...
By Shuhong Zheng, Zhipeng Bao, Ruoyu Zhao, Martial Hebert, Yu-Xiong Wang
arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv:2509.01167v3 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current...
By Hyunjong Ok, Jaeho Lee
arXiv:2511.07260v3 Announce Type: replace-cross
Abstract: Ad hoc teamwork (AHT) requires agents to collaborate with previously unseen teammates, which is crucial for many real-world applications. The...
By Hohei Chan, Xinzhi Zhang, Antao Xiang, Weinan Zhang, Mengchen Zhao
arXiv:2603.01590v2 Announce Type: replace-cross
Abstract: Content-driven platforms such as Xiaohongshu often leverage click-through rate (CTR) prediction models for recommendation. However, these mod...
By Yubin Zhang, Haiming Xu, Guillaume Salha-Galvan, Ruiyan Han, Feiyang Xiao, Yanhua Huang, Li Lin, Yang Luo, Yao Hu
arXiv:2608.29601v2 Announce Type: replace-cross
Abstract: We present $N_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale mul...
By NeoteAI Team, Fudan TEAI Team
arXiv:2609.22210v1 Announce Type: new
Abstract: SALSA (Semi-Autonomous Literature Summarization Assistant) is an open- source, human-in-the-loop platform for extracting structured scientific datasets...
By William Schertzer, Sonakshi Gupta, Rampi Ramprasad