arXiv:2607. 12364v1 Announce Type: cross Abstract: EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning.
By Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik
arXiv:2511. 09483v3 Announce Type: replace Abstract: While multimodal large language models can describe visual content, their ability to generate executable procedures remains underexplored.
By Peiyu Li, Xiaobao Huang, Ting Hua, Nitesh V. Chawla
arXiv:2607. 12290v1 Announce Type: cross Abstract: Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation.
By Chun-Yi Kuan, Hung-yi Lee
arXiv:2607. 12468v1 Announce Type: cross Abstract: We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time.
By Shuming Fang, Shuifei Zeng
arXiv:2606. 20101v3 Announce Type: replace-cross Abstract: Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content.
By Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang, Shubin Zhang, Zhenbo Li, Jean-Yves Guillemaut, Wenwu Wang
arXiv:2607. 13013v1 Announce Type: new Abstract: Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time.
By Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal
arXiv:2602. 15892v2 Announce Type: replace-cross Abstract: Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition.
By Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo
arXiv:2607. 12375v1 Announce Type: cross Abstract: Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability.
By Jinjian Wu, Jiaqi Tang, Wei Wei, Yingying Yan, Jianmin Chen, Botong Geng, Lei Zhang, Qifeng Chen
arXiv:2607. 12711v1 Announce Type: new Abstract: Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based puzzles.
By Pedro Orvalho, Guillem Aleny\`a, Felip Many\`a
arXiv:2607. 12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition.
By Aleksei Bakin, Andrey V. Savchenko
arXiv:2607. 12281v1 Announce Type: cross Abstract: Transformer-style architectures are increasingly adopted for industrial recommendation systems, yet they inherit a design premise misaligned with the task: generative models rely on per-token autoregressive prediction, which justifies maintaining large intermediate tensors that scale with sequence length.
By Siqi Wang, Xianjie Chen, Shaofeng Deng, Albert Chen, Romil Shah, Jiawei Huang, Zhaoqin Wang, Zhang Zhang, Yiqun Liu, Meilei Jiang, Anish Dubey, Moyan Mei, Tongxin Wang, Nathan Berrebbi, Misael Manjarres, Armand Sauzay, Shardul Kothapalli, Aryaman Vinchhi, Kevin Johnstone, Juheon Lee, Gufan Yin, Ziheng Huang, Justin Lin, Mert Terzihan, Yilin Qi, Cynthia Yang, Colin Peppler, Qi Ding, Ruohan Sun, Ge Song, Litao Deng, Parichay Kapoor, Matt Ma, Huihui Cheng, Jiyuan Zhang, Yanli Zhao, Yiping Han, Fangqiu Han, Ning Yao, Arun Singh, Jordan Edwards, Zhengyu Su, Abhishek Kumar, Guangdeng Liao, Ankit Asthana
arXiv:2509. 22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect.
By Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
arXiv:2602. 18548v2 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inconsistent datasets, toolchains, and evaluation protocols.
By Qiao Xu, Yipeng Yu, Chengxiao Feng, Xu Liu
arXiv:2607. 12678v1 Announce Type: cross Abstract: Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geometric and semantic cues for intelligent design understanding.
By Yan Gong, Bohao Li, Bowen Du, Junchen Ye
arXiv:2603. 16307v2 Announce Type: replace Abstract: Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex scenes and constraints and make reliable decisions.
By Ming Yang, Zhi Zhou, Shi-Yu Tian, Kun-Yang Yu, Lan-Zhe Guo, Yu-Feng Li
arXiv:2607. 12982v1 Announce Type: new Abstract: Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples.
By Ruoran Xu, Wending Gao, Qiufeng Wang
arXiv:2512. 15748v2 Announce Type: replace Abstract: Visual Species Recognition (VSR) is a fundamental task in scientific disciplines that require species-level identification, including ecology, palynology, evolutionary biology, systematics, and phylogenetics.
By Tian Liu, Anwesha Basu, James Caverlee, Shu Kong
arXiv:2607. 12833v1 Announce Type: cross Abstract: Circular data, representing angles or directions, are frequently encountered in computer vision, biology, geology, and meteorology.
By Rajdeep Pathak, Archi Roy, Tanujit Chakraborty
arXiv:2607. 12065v1 Announce Type: cross Abstract: While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions.
By Robel Mamo, Rajitha de Silva, Grzegorz Cielniak, Taeyeong Choi
arXiv:2209. 14125v3 Announce Type: replace-cross Abstract: Diffusion models have proven to be a flexible and effective framework for modelling probability distributions on finite-dimensional spaces.
By Angus Phillips, Thomas Seror, Michael Hutchinson, Valentin De Bortoli, Arnaud Doucet, Emile Mathieu