Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv AI
3d ago

Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration

The paper introduces Quantizer‑Aligned Recalibration (QuAR), a single‑pass test‑time adaptation technique for quantized vision transformers that does not require backpropagation or parameter updates. QuAR recalibrates activations at the input of frozen quantizers by aligning per‑channel statistics with the source calibration, thereby correcting the distorted code distribution caused by distribution shift. On ImageNet‑C, QuAR outperforms state‑of‑the‑art backprop‑free methods across 3‑, 4‑, 6‑, and 8‑bit precisions, achieving higher accuracy, lower latency, and minimal memory overhead while maintaining performance across diverse shift scenarios.

By Hyeongheon Cha, Young D. Kwon, Sung-Ju Lee
arXiv AI
3d ago

Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

The paper presents a new multi‑turn benchmark of 423 conversations with 1,661 labeled turn‑states to study when language models should refrain from answering. It shows that probes for unanswerability transfer well across datasets that share the same underlying signal, but fail to generalise to other forms of epistemic uncertainty. While a calibrated probe can identify underspecified turns more accurately than chance, it does not consistently improve overall generation quality compared to standard methods.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya
arXiv AI
3d ago

Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment

arXiv:2610.08482v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited,...

By Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, Xiaojuan Li, Mingrui Yang
arXiv AI
3d ago

Language-model ratings of depression reflect the rater more than the patient

The study examined how language‑model raters assess depression using the Patient Health Questionnaire across 880 raters and 189 interviews. Model choice accounted for 30% of symptom‑score variance, while stable participant differences explained 10.5%. Even raters with similar overall accuracy (AUC ≥ 0.70) disagreed on screening decisions for 40% of participants, and only recalibration with labeled data improved agreement and accuracy modestly.

By Baihan Lin
arXiv AI
3d ago

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

The paper argues that probe scores lack intrinsic meaning and should be interpreted relative to two reference points: a floor (what simple inputs predict) and a ceiling (what the full input predicts). The difference, called headroom, indicates the range where a probe can reveal that a model computes beyond what the input already provides. Experiments on transformers and real models show that headroom can vanish when the target no longer depends on hidden variables or when the input no longer reveals them, and that some previously claimed representations are largely explained by the input text alone.

By Pranjal Garg
arXiv AI
3d ago

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

arXiv:2605.24154v2 Announce Type: replace Abstract: Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users an...

By Qitao Tan, Xiaoying Song, Arman Akbari, Arash Akbari, Yanzhi Wang, Xiaoming Zhai, Lingzi Hong, Zhen Xiang, Jin Lu, Geng Yuan