arXiv:2609.22628v1 Announce Type: new
Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We iso...
By Nikhil Reddy Pottanigari, Sepideh Kharaghani, Saverio Vadacchino, Alejandro Posada, Ying Zhang
arXiv:2609.22790v1 Announce Type: new
Abstract: Undergraduate programmes in intelligent medical engineering are expanding, yet laboratory curricula lag behind the multimodal, long-tailed, and distrib...
By Dongjing Shan, Yamei Luo, Jin Li, Yong Luo
arXiv:2609.22910v1 Announce Type: new
Abstract: Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the mod...
By Kunyu Peng, Junming Liu, Ruiqi He, Qingzhuo Wang, Jianzhong Qi, Xianhui Liu
arXiv:2609.23064v1 Announce Type: new
Abstract: Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical sys...
By Qiang Chen, Hao Guo, Huatai Zhu, Tairan Huang, Yichao Cao, Hongyan Xu, Keke Huang, Haifeng Li, Yi Chen, Xiu Su
arXiv:2609.23130v1 Announce Type: new
Abstract: Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state,...
By Twinkll Sisodia
arXiv:2609.23860v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the features of these ViTs are ultimately consum...
By Tianyou Jiang
arXiv:2609.24092v1 Announce Type: new
Abstract: Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page....
By Jeremy Cerwin Wang, Wai Kit Wong, Jeff Kai Tai Tang
arXiv:2609.24362v1 Announce Type: new
Abstract: Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language m...
By Hexiong Yang, Mingrui Chen, Jie Cao, Ran He
arXiv:2609.24677v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cann...
By Jie Gong, Maowei Jiang, Zhiwei Liu, Yankai Chen, Guojun Xiong, Xue Liu, Min Peng, Qianqian Xie, Sophia Ananiadou
arXiv:2609.22538v1 Announce Type: cross
Abstract: Large language model (LLM) planners can decompose natural-language instructions and select reusable robot skills, but choosing the correct skill does...
By Ajay Vikram Periasami, Xinyuan Luo, Haoyu Li, Xianyi Cheng
arXiv:2609.23407v1 Announce Type: cross
Abstract: Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for em...
By Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong
arXiv:2609.24094v1 Announce Type: cross
Abstract: Visual analytics (VA) enables sensemaking through interactive visualization, but effective analysis often requires experts to translate high-level in...
By Yutong Chen, Zhike Tang, Zhihao Mai, Zhihao Shuai, Danli Luo, Jing Xu, Weikai Yang
arXiv:2605.16638v2 Announce Type: replace
Abstract: Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this pa...
By Jianpeng Cheng, Xian Wu, Jiangfan Zhang, Wentao Bao, Chaitanya Ahuja, Shlok Kumar Mishra, Xuanming Cui, Hanchao Yu, Yang Gao, Fan Xia, Haixing Dai, Frankie Yuan, Zihao Wang, Xiaobing Chen, Qi Guo, Shaodan Zhai, Aashu Singh, Xiangjun Fan, Jun Xiao
arXiv:2510.26915v2 Announce Type: replace-cross
Abstract: While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large...
By Zachary Ravichandran, Fernando Cladera, Ankit Prabhu, Jason Hughes, Carlos Nieto-Granda, Varun Murali, Camillo Taylor, George J. Pappas, Vijay Kumar
arXiv:2603.03768v2 Announce Type: replace-cross
Abstract: Full-stack human-robot collaboration (HRC) can become brittle when replacing a planner, partner model, coordination policy, or controller cha...
By Hao Zhang, Yisen Li, Ruize Geng, Yves Tseng, Yaru Niu, Ding Zhao, H. Eric Tseng
arXiv:2609.25131v1 Announce Type: new
Abstract: Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also expo...
By Tung Sum Thomas Kwok, Yidong Ouyang, Yingjia Wan, Ying Nian Wu, Zhijiang Guo, Oscar Leong
arXiv:2609.25146v1 Announce Type: new
Abstract: Continual learning, the ability to learn from sequential experience while retaining and adapting prior knowledge, is central to intelligent systems ope...
By Hongwei Yan, Kanglei Zhou, Qi Cheng, Weiyi Dong, Chunyan Lan, Guanglong Sun, Jun Zhou, Qian Li, Yi Zhong, Liyuan Wang
arXiv:2609.25152v1 Announce Type: new
Abstract: Deep Imbalanced Regression (DIR) addresses a common failure mode of regression models: target distributions are highly non-uniform, causing models to p...
By Noah C. Puetz, Jens U. Brandt, Marc Hilbert, Elena Raponi, Thomas B\"ack, Thomas Bartz-Beielstein
arXiv:2609.25809v1 Announce Type: new
Abstract: Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly man...
By Yuanteng Chen, Qiwei Lai, Chen Tianqi, Peisong Wang, Yuantian Shao, Nanxin Zeng, Zhilei Liu, Chuangyi Li, Jing Liu, Jian Cheng
arXiv:2609.25411v1 Announce Type: cross
Abstract: Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpola...
By Biel Tura Vecino, Yoach Lacombe, Julian Weber, Zbigniew {\L}atka, Haitong Zhang, Logan Hart, Eren G\"olge