arXiv:2608.16697v2 Announce Type: replace
Abstract: Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heter...
By Aniri, Chen Yilin, Jinhe Bi, Zengjie Jin, Yujun Wang, Yijun Tian, Volker Tresp, Fei Shen, Tat-Seng Chua, Yunpu Ma
arXiv:2609.38368v1 Announce Type: new
Abstract: Vision-language models (VLMs) increasingly reason over visual evidence that is cropped, segmented, retrieved, or revealed over time. Yet most VQA bench...
By L. D. M. S. Sai Teja, Ufaq Khan, N. Siva Gopala Krishna, Satyajit Tourani, Ashshak Sharifdeen, Fida Mohammad Thoker, Bernard Ghanem, Muhammad Haris Khan
arXiv:2609.38620v1 Announce Type: new
Abstract: Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-...
By Hanwen Cao, Wenqiang Wu, Kuang-Ting Tu, Mathias Otnes, Jeffrey Delmerico, Rui Wang, Yulun Tian, Nikolay Atanasov
arXiv:2609.38823v1 Announce Type: new
Abstract: Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly becau...
By Xudong Tan, Peng Ye, Ming Xie, Chenyu Huang, Yaoxin Yang, Jiayuan Fan, Tao Chen
arXiv:2609.39021v1 Announce Type: new
Abstract: Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly...
By Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang
arXiv:2609.39235v1 Announce Type: cross
Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
By Ali Alrasheed, Basim Azam, Naveed Akhtar
arXiv:2609.31665v2 Announce Type: replace
Abstract: Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce la...
By Mahsa Mohammadi, Zeyu Fu, Sareh Rowlands
arXiv:2609.39964v1 Announce Type: new
Abstract: Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. D...
By Yijie Bian, Kai Zhang, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief
arXiv:2609.38523v1 Announce Type: cross
Abstract: Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle co...
By Dong Shu, Yanguang Liu, Huopu Zhang, Saisai Hu, Haiyan Zhao, Hekun Huang, Mengnan Du
ChartDensity-Bench is a benchmark designed to evaluate multimodal large language models (MLLMs) on their ability to reconstruct structured numerical data from scientific charts that vary in visual density. The benchmark uses charts paired with source-level ground-truth data and systematically changes the number of simultaneously presented charts (k = 1, 3, 6, 9) to assess how density affects reconstruction performance. A multi‑dimensional evaluation framework measures structural reliability, reconstruction completeness, parseability, and numerical fidelity, revealing that numerical reconstruction generally worsens as visual density increases, with varying degrees of degradation across different models.
By Xinhe Wu, Yadong Jin
The paper introduces MEQ, a mutual feedback architecture that iteratively refines two multimodal inputs into coupled embeddings, each embedding incorporating information from the other. By continuously exchanging information between the modalities, the model converges to a fixed point that improves representation quality. Experiments on classification and visual grounding tasks show that MEQ achieves competitive or superior performance compared to concatenation-based baselines, and qualitatively enhances visual grounding when paired with complementary modalities.
By Ho-min Park, Byungkon Kang
arXiv:2609.40361v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based o...
By Tian Xia, Minghao Liu, Yiqing Liang, Laixi Shi, Jiayun Wang
The TutlAit v1 dataset is a crowdsourced corpus of Moroccan Tamazight speech paired with Modern Standard Arabic transcriptions and explicit regional accent labels. It contains 13,384 audio files (≈20.9 hours) collected via a web application, with volunteers contributing through text‑to‑audio and audio‑to‑text workflows, and includes additional segments from freely available media. The dataset covers Atlas, Souss, Rif, and Kabyle varieties, making it a valuable resource for speech recognition, translation, and accent identification in an under‑resourced language.
By Mohamed-Amine Chadi, Ezzahra Ait El Arbi, Ismail Khayoub, Aymane Fadili, Yassine Ennhili, Jadjigua Bouali, Hanane Inhid, Mohammed Ameksa, Hajar Mousannif
arXiv:2609.38925v1 Announce Type: new
Abstract: Multimodal federated learning (MFL) has emerged as a pivotal paradigm for leveraging distributed data to enhance model performance. However, existing m...
By Tianchi Liao Tianchi_Liao, Lele Fu, Sheng Huang, Qing Hu, Hong-Ning Dai, Chuan Chen
arXiv:2609.39047v1 Announce Type: new
Abstract: Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, ye...
By Zhihang Wu, Zhongqi Wang, Jie Zhang, Fengming Gu, Shiguang Shan, Xilin Chen
arXiv:2609.39588v1 Announce Type: new
Abstract: We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical la...
By Aravindh Mahendran, Michael King, Matthew Koichi Grimes, Antoine Yang, Tyler Zhu, Joseph Heyward, Tengda Han, Shiry Ginosar, Chen Sun, Dima Damen, Simon Osindero, Noah Snavely, Simon Lynen, Jo\~ao Carreira, Viorica P\u{a}tr\u{a}ucean
GroundAnything is a 4‑B parameter grounding foundation model that combines autoregressive and diffusion approaches to achieve fast parallel decoding while maintaining precise visual grounding. By treating grounding as visual evidence extraction and using blockwise denoising, it allows spatial hypotheses to be generated in parallel and refined iteratively. The model outperforms existing state‑of‑the‑art methods on 30 grounding benchmarks, achieving 72.42% accuracy with its autoregressive variant and 61.75% with entropy‑guided decoding, while also offering significant speedups through optional self‑speculative decoding.
By Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou, Ping Luo, Shiyu Huang
arXiv:2609.39836v1 Announce Type: new
Abstract: Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natura...
By Simone Ricci, Niccol\`o Biondi, Federico Pernici
BayesNDE is a neural density estimator that uses Bayesian generative modeling to estimate densities without relying on invertible networks or Jacobian-determinant calculations. It constructs an adaptive proposal for each observation by inferring a sample-specific latent posterior, and then applies bridge sampling to combine proposal samples with separate posterior samples for density estimation. Experiments on synthetic datasets show improved density estimation and structure recovery, while real-world applications demonstrate better anomaly detection.
By Chenglin Li, Qiao Liu
CollageAttack is a black‑box jailbreak that exploits cross‑modal alignment flaws in text‑to‑image models by shifting harmful semantics into the image plane. It combines context‑relevant scenes, scene‑grounded textual carriers, and spatially distributed text fragments to produce images that reveal hidden harmful meaning. Experiments on both open‑weight and commercial models show success rates up to 86.0%, outperforming the strongest baseline by 18.5 percentage points and consistently generating more harmful outputs while preserving the source intent.
By Zhiyi Mou, Yao Lu, Wangze Ni, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou, Kui Ren