arXiv:2610.07832v1 Announce Type: cross
Abstract: Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, e...
By Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang
arXiv:2610.07862v1 Announce Type: cross
Abstract: A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce G...
By Bin Cao, Huichi Zhou, Runyu Yang, Jingsong Li, Shuchen Sun, Yan Song, Hanyu Gao, Zhongwei Yu, Tong-Yi Zhang, Jun Wang
arXiv:2610.07863v1 Announce Type: cross
Abstract: Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with s...
By Yupeng Su, Jiayi Tian, Zheng Zhang, Souvik Kundu
arXiv:2610.07913v1 Announce Type: cross
Abstract: Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whol...
By Shrihari Dumbre, Bikash Santra
arXiv:2610.07940v1 Announce Type: cross
Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key...
By Yuhan Chen, Siyuan Zhang, Nan Wang, Feiyang Kang, Ruoxi Jia
arXiv:2610.07996v1 Announce Type: cross
Abstract: Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any par...
By Yacine Benihaddadene, Milan Bhan, Eliot Dugelay, Mohammed Jawhar, Benjamin Wong, Nicolas Chesneau, Duong Nguyen
arXiv:2610.08093v1 Announce Type: cross
Abstract: Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-q...
By Chuan Li, Chengyu Wang, Cen Chen, Ye Lyu, Mingyuan Fan, Ming Gao
arXiv:2610.08161v1 Announce Type: cross
Abstract: Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedC...
By Daniel Varab, Victor Petr\'en Bach Hansen, Asbj{\o}rn W. Helge, Kevin Pelgrims, Mathias Baltzersen, Adrian Young-San Roessler, Vanessa Klungtvedt, Maximilian Brand, Lasse Krogsb{\o}ll, Henrik Cullen, Lars Maal{\o}e
arXiv:2610.07887v1 Announce Type: cross
Abstract: Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand abo...
By Chufan Shi, Cheng Yang, Tiannuo Yang, Isadora White, Yiwei Chen, Taylor Berg-Kirkpatrick, Xuezhe Ma
arXiv:2610.08388v1 Announce Type: cross
Abstract: Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on kno...
By Yang Hong, Yajun Yang, Xin Wang, Liping Jing, Qinghua Hu
arXiv:2610.08401v1 Announce Type: cross
Abstract: While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information...
By Seulgi Kim, Zhixiong Zhang, Xinwei Zhang, Jie Ling, Ronn Shaw
arXiv:2610.08528v1 Announce Type: cross
Abstract: Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image featu...
By Asim Khan, Samee Ullah Khan, Dwarikanath Mahapatra
arXiv:2610.08538v1 Announce Type: cross
Abstract: Probabilistic load forecasting has been widely studied for power-system operation and planning, but customer- and transformer-level forecasting intro...
By Haoran Li, Zhe Cheng, Yang Weng
arXiv:2610.08559v1 Announce Type: cross
Abstract: Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports...
By Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban, Stuart Burrell
arXiv:2610.08571v1 Announce Type: cross
Abstract: Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a be...
By Niveen O. Jaffal, Ahmet Yuksel, David Mohaisen
arXiv:2610.08577v1 Announce Type: cross
Abstract: While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their...
By Hwiyeong Lee, Hyelim Lim, Ingyu Bang, Hoki Kim, Taeuk Kim
arXiv:2610.08642v1 Announce Type: cross
Abstract: Contact with contaminated objects can spread hazards through a household robot's grippers, tools, and shared surfaces, while new contacts can make an...
By Yurun Chen, Josh Qixuan Sun, Jason Qin, Chengtai Li, Tianyi Wang, Mark Crowley, Wentao Zhu
arXiv:2610.08670v1 Announce Type: cross
Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does...
By Orion Reblitz-Richardson
arXiv:2610.08678v1 Announce Type: cross
Abstract: Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model,...
By Yichi Zhang, Zhiqi Wang, Neil Gong, Yuchen Yang
arXiv:2610.08781v1 Announce Type: cross
Abstract: Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training...
By Ziyu Chen, Yilun Zhao, Jiashuo Sun, Yiling Ma, Manasi Patwardhan, Arman Cohan