The paper introduces Quantizer‑Aligned Recalibration (QuAR), a single‑pass test‑time adaptation technique for quantized vision transformers that does not require backpropagation or parameter updates. QuAR recalibrates activations at the input of frozen quantizers by aligning per‑channel statistics with the source calibration, thereby correcting the distorted code distribution caused by distribution shift. On ImageNet‑C, QuAR outperforms state‑of‑the‑art backprop‑free methods across 3‑, 4‑, 6‑, and 8‑bit precisions, achieving higher accuracy, lower latency, and minimal memory overhead while maintaining performance across diverse shift scenarios.
By Hyeongheon Cha, Young D. Kwon, Sung-Ju Lee
arXiv:2610.08388v1 Announce Type: cross
Abstract: Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on kno...
By Yang Hong, Yajun Yang, Xin Wang, Liping Jing, Qinghua Hu
arXiv:2610.08405v1 Announce Type: cross
Abstract: Large Language Models have emerged as promising tools for software vulnerability analysis, but their effectiveness depends heavily on prompt design....
By Mandana Ghadamian, David Mohaisen
The paper presents a new multi‑turn benchmark of 423 conversations with 1,661 labeled turn‑states to study when language models should refrain from answering. It shows that probes for unanswerability transfer well across datasets that share the same underlying signal, but fail to generalise to other forms of epistemic uncertainty. While a calibrated probe can identify underspecified turns more accurately than chance, it does not consistently improve overall generation quality compared to standard methods.
By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya
arXiv:2610.08482v1 Announce Type: cross
Abstract: Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited,...
By Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, Xiaojuan Li, Mingrui Yang
The study examined how language‑model raters assess depression using the Patient Health Questionnaire across 880 raters and 189 interviews. Model choice accounted for 30% of symptom‑score variance, while stable participant differences explained 10.5%. Even raters with similar overall accuracy (AUC ≥ 0.70) disagreed on screening decisions for 40% of participants, and only recalibration with labeled data improved agreement and accuracy modestly.
By Baihan Lin
arXiv:2610.08513v1 Announce Type: cross
Abstract: LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human in...
By Dennis Fucci, Andrea Bacciu, Dong Liu, Weronika {\L}ajewska, Saab Mansour
The paper argues that probe scores lack intrinsic meaning and should be interpreted relative to two reference points: a floor (what simple inputs predict) and a ceiling (what the full input predicts). The difference, called headroom, indicates the range where a probe can reveal that a model computes beyond what the input already provides. Experiments on transformers and real models show that headroom can vanish when the target no longer depends on hidden variables or when the input no longer reveals them, and that some previously claimed representations are largely explained by the input text alone.
By Pranjal Garg
arXiv:2610.08559v1 Announce Type: cross
Abstract: Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports...
By Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban, Stuart Burrell
arXiv:2610.08622v1 Announce Type: cross
Abstract: System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troub...
By Sayan Sinha, Vipul Harsh, B. Aditya Prakash, Vyas Sekar, Hui Zhang
arXiv:2610.08659v1 Announce Type: cross
Abstract: Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-bas...
By Suxin Ji, Hungtao Wan, Mingjun Liu, An Zhang
arXiv:2610.08668v1 Announce Type: cross
Abstract: Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Pri...
By Suxin Ji, Hungtao Wan, Shaoxuan Chen, An Zhang
arXiv:2610.08669v1 Announce Type: cross
Abstract: On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficien...
By Mehmet Emre Akbulut, Johannes Geier, Ulf Schlichtmann
arXiv:2610.08670v1 Announce Type: cross
Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does...
By Orion Reblitz-Richardson
arXiv:2507.07426v4 Announce Type: replace
Abstract: Recent advances in large language models have demonstrated considerable potential in scientific domains such as drug repositioning. However, their...
By Zerui Yang, Yuwei Wan, Siyu Yan, Yudai Matsuda, Tong Xie, Linqi Song
arXiv:2508.16821v2 Announce Type: replace
Abstract: We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinfo...
By Sam Earle, Graham Todd, Yuchen Li, Ahmed Khalifa, Muhammad Umair Nasir, Zehua Jiang, Andrzej Banburski-Fahey, Julian Togelius
arXiv:2511.20471v3 Announce Type: replace
Abstract: Recent advances in Large Language Model (LLM) reasoning have improved conventional problem solving, but creative reasoning remains comparatively un...
By Yuto Suzuki, Farnoush Banaei-Kashani
arXiv:2605.10791v2 Announce Type: replace
Abstract: Knowledge Graph Question Answering (KGQA) aims to answer user questions by reasoning over Knowledge Graphs (KGs). Recent methods use supervision de...
By Shengxiang Gao, Chao Lei, Jey Han Lau, Linhao Luo, Jianzhong Qi
arXiv:2605.12922v2 Announce Type: replace
Abstract: Large language models can follow complex instructions in a single turn, yet over long multi-turn interactions they often lose the thread of instruc...
By Vardhan Dongre, Joseph Hsieh, Viet Dac Lai, Seunghyun Yoon, Trung Bui, Dilek Hakkani-T\"ur
arXiv:2605.24154v2 Announce Type: replace
Abstract: Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users an...
By Qitao Tan, Xiaoying Song, Arman Akbari, Arash Akbari, Yanzhi Wang, Xiaoming Zhai, Lingzi Hong, Zhen Xiang, Jin Lu, Geng Yuan