arXiv:2610.01166v1 Announce Type: cross
Abstract: Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ven...
By Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang
arXiv:2601.09879v2 Announce Type: replace-cross
Abstract: Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report gen...
By Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
arXiv:2609.38764v1 Announce Type: new
Abstract: Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract,...
By Zachary Shinnick, Hemanth Saratchandran, Damien Teney, Anton van den Hengel
arXiv:2609.39880v1 Announce Type: cross
Abstract: Passwords remain the dominant online authentication mechanism, and understanding how humans choose them is essential for defensive strength estimatio...
By Rajneesh Anand, Neeraj Lakshmanan, Masoud Yari
arXiv:2609.40016v1 Announce Type: cross
Abstract: Exact incremental BPE maintains the canonical tokenization state after every appended byte. The recent algorithm of Jiang and Gong (2026) does this i...
By Harshit Verma, Rex Ying
arXiv:2610.00573v1 Announce Type: new
Abstract: Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant fram...
By Haifeng Huang, Biyin Xu, Chunsheng Xin, Yang Li
arXiv:2610.00686v1 Announce Type: new
Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene...
By Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss
arXiv:2610.00757v1 Announce Type: new
Abstract: Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may re...
By Haowen Guan, Shengzhi Li, Shichao Pei
arXiv:2610.00881v1 Announce Type: new
Abstract: Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation...
By Ozge Mercanoglu Sincan, Anton Pelykh, Edward Fish, Harry Walsh, JianHe Low, Karahan Sahin, Oline Ranum, Sobhan Asasi, Steven Emery, Richard Bowden
arXiv:2610.01180v1 Announce Type: new
Abstract: Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently str...
By Yuliang Cai, Mohammad Rostami, Jesse Thomason
arXiv:2610.01942v1 Announce Type: new
Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of...
By Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
arXiv:2606.21562v2 Announce Type: replace
Abstract: Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming v...
By Philippe Weinzaepfel, Christian Wolf, Mert B\"ulent Sariyildiz, Guillaume Bono, Gianluca Monaci
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
The paper introduces a method to purify LoRA-tuned large language models (LLMs) against backdoor attacks without relying on trigger knowledge, clean references, or retraining. By extracting high‑fidelity backdoor directions and projecting LoRA updates onto orthogonal null spaces in input and output channels, the approach reduces attack success rates from nearly 100% to under 10%. Experiments demonstrate that this null‑space projection preserves both the base model’s general capabilities and the new downstream skills learned through the adapter across various tasks.
By Jianwei Li, Jung-Eun Kim
arXiv:2610.00899v1 Announce Type: cross
Abstract: Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with st...
By Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae
arXiv:2404.11624v3 Announce Type: replace-cross
Abstract: We introduce Token Space, a categorical framework for AI computations based on explicit structural records. Five theses guide it: object inte...
By Wuming Pan
arXiv:2609.40245v2 Announce Type: cross
Abstract: Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-...
By Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas
arXiv:2609.39222v1 Announce Type: new
Abstract: High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff be...
By Xu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou
ChartDensity-Bench is a benchmark designed to evaluate multimodal large language models (MLLMs) on their ability to reconstruct structured numerical data from scientific charts that vary in visual density. The benchmark uses charts paired with source-level ground-truth data and systematically changes the number of simultaneously presented charts (k = 1, 3, 6, 9) to assess how density affects reconstruction performance. A multi‑dimensional evaluation framework measures structural reliability, reconstruction completeness, parseability, and numerical fidelity, revealing that numerical reconstruction generally worsens as visual density increases, with varying degrees of degradation across different models.
By Xinhe Wu, Yadong Jin
arXiv:2609.39321v1 Announce Type: cross
Abstract: Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks....
By Rajat Ghosh, Vaishnavi Bhargava, Henry Wong, Aryan Singhal, Debojyoti Dutta