arXiv:2610.07940v1 Announce Type: cross
Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key...
By Yuhan Chen, Siyuan Zhang, Nan Wang, Feiyang Kang, Ruoxi Jia
VisionWeave introduces elastic visual representation weaving, a native capability for multimodal large language models that learns where and at what granularity to encode visual information. The method combines a gated spatial pooler for coarse representations with a granularity router that allocates content‑adaptive token usage, trained end‑to‑end on large‑scale data. Experiments on Qwen3.5‑4B and Qwen3.8‑27B show that VisionWeave can save 43.0% of tokens while preserving 98.9% of performance across eight benchmarks, and delivers significant throughput gains and latency reductions when deployed on the SGLang serving engine.
By Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, Xike Xie
arXiv:2610.08093v1 Announce Type: cross
Abstract: Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-q...
By Chuan Li, Chengyu Wang, Cen Chen, Ye Lyu, Mingyuan Fan, Ming Gao
The paper introduces a new evaluation setting called penalty‑framed no‑valid‑option MCQA, where multiple‑choice questions may contain no correct answer. By removing the correct option from the MMLU‑Pro mathematics subset and allowing models to either pick an option or abstain, the authors penalize forced‑choice responses that are invalid. Experiments reveal that even models with high standard MCQA accuracy can still produce invalid forced‑choice answers, indicating that traditional accuracy metrics miss an important aspect of model reliability.
By Jinhyeok Kim, Hye-Young Jung
arXiv:2610.08155v1 Announce Type: cross
Abstract: Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by en...
By Zihan Zhou, Xinzhe Hu, Hanxu Yang, Liangjian Wen, Zhao Kang
arXiv:2610.08183v1 Announce Type: cross
Abstract: Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performa...
By Nanhe Chen, Runqiu Yang, Jiawei Tang, Sichao Liu, Yuquan Wang
arXiv:2610.08268v1 Announce Type: cross
Abstract: Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models...
By Jingpo Xu, Paul Joe Maliakel, Ivona Brandic, Shashikant Ilager
arXiv:2610.08331v1 Announce Type: cross
Abstract: The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in ex...
By Heyam Bin Jahlan Areej Alhothali Abeer Alhothali
The paper introduces Quantizer‑Aligned Recalibration (QuAR), a single‑pass test‑time adaptation technique for quantized vision transformers that does not require backpropagation or parameter updates. QuAR recalibrates activations at the input of frozen quantizers by aligning per‑channel statistics with the source calibration, thereby correcting the distorted code distribution caused by distribution shift. On ImageNet‑C, QuAR outperforms state‑of‑the‑art backprop‑free methods across 3‑, 4‑, 6‑, and 8‑bit precisions, achieving higher accuracy, lower latency, and minimal memory overhead while maintaining performance across diverse shift scenarios.
By Hyeongheon Cha, Young D. Kwon, Sung-Ju Lee
arXiv:2610.08388v1 Announce Type: cross
Abstract: Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on kno...
By Yang Hong, Yajun Yang, Xin Wang, Liping Jing, Qinghua Hu
arXiv:2610.08405v1 Announce Type: cross
Abstract: Large Language Models have emerged as promising tools for software vulnerability analysis, but their effectiveness depends heavily on prompt design....
By Mandana Ghadamian, David Mohaisen
The paper presents a new multi‑turn benchmark of 423 conversations with 1,661 labeled turn‑states to study when language models should refrain from answering. It shows that probes for unanswerability transfer well across datasets that share the same underlying signal, but fail to generalise to other forms of epistemic uncertainty. While a calibrated probe can identify underspecified turns more accurately than chance, it does not consistently improve overall generation quality compared to standard methods.
By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya
arXiv:2610.08482v1 Announce Type: cross
Abstract: Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited,...
By Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, Xiaojuan Li, Mingrui Yang
The study examined how language‑model raters assess depression using the Patient Health Questionnaire across 880 raters and 189 interviews. Model choice accounted for 30% of symptom‑score variance, while stable participant differences explained 10.5%. Even raters with similar overall accuracy (AUC ≥ 0.70) disagreed on screening decisions for 40% of participants, and only recalibration with labeled data improved agreement and accuracy modestly.
By Baihan Lin
arXiv:2610.08513v1 Announce Type: cross
Abstract: LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human in...
By Dennis Fucci, Andrea Bacciu, Dong Liu, Weronika {\L}ajewska, Saab Mansour
The paper argues that probe scores lack intrinsic meaning and should be interpreted relative to two reference points: a floor (what simple inputs predict) and a ceiling (what the full input predicts). The difference, called headroom, indicates the range where a probe can reveal that a model computes beyond what the input already provides. Experiments on transformers and real models show that headroom can vanish when the target no longer depends on hidden variables or when the input no longer reveals them, and that some previously claimed representations are largely explained by the input text alone.
By Pranjal Garg
arXiv:2610.08559v1 Announce Type: cross
Abstract: Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports...
By Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban, Stuart Burrell
arXiv:2610.08622v1 Announce Type: cross
Abstract: System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troub...
By Sayan Sinha, Vipul Harsh, B. Aditya Prakash, Vyas Sekar, Hui Zhang
arXiv:2610.08659v1 Announce Type: cross
Abstract: Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-bas...
By Suxin Ji, Hungtao Wan, Mingjun Liu, An Zhang
arXiv:2610.08668v1 Announce Type: cross
Abstract: Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Pri...
By Suxin Ji, Hungtao Wan, Shaoxuan Chen, An Zhang