Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv AI
2d ago

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

VisionWeave introduces elastic visual representation weaving, a native capability for multimodal large language models that learns where and at what granularity to encode visual information. The method combines a gated spatial pooler for coarse representations with a granularity router that allocates content‑adaptive token usage, trained end‑to‑end on large‑scale data. Experiments on Qwen3.5‑4B and Qwen3.8‑27B show that VisionWeave can save 43.0% of tokens while preserving 98.9% of performance across eight benchmarks, and delivers significant throughput gains and latency reductions when deployed on the SGLang serving engine.

By Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, Xike Xie
arXiv AI
2d ago

Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices

The paper introduces a new evaluation setting called penalty‑framed no‑valid‑option MCQA, where multiple‑choice questions may contain no correct answer. By removing the correct option from the MMLU‑Pro mathematics subset and allowing models to either pick an option or abstain, the authors penalize forced‑choice responses that are invalid. Experiments reveal that even models with high standard MCQA accuracy can still produce invalid forced‑choice answers, indicating that traditional accuracy metrics miss an important aspect of model reliability.

By Jinhyeok Kim, Hye-Young Jung
arXiv AI
2d ago

Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration

The paper introduces Quantizer‑Aligned Recalibration (QuAR), a single‑pass test‑time adaptation technique for quantized vision transformers that does not require backpropagation or parameter updates. QuAR recalibrates activations at the input of frozen quantizers by aligning per‑channel statistics with the source calibration, thereby correcting the distorted code distribution caused by distribution shift. On ImageNet‑C, QuAR outperforms state‑of‑the‑art backprop‑free methods across 3‑, 4‑, 6‑, and 8‑bit precisions, achieving higher accuracy, lower latency, and minimal memory overhead while maintaining performance across diverse shift scenarios.

By Hyeongheon Cha, Young D. Kwon, Sung-Ju Lee
arXiv AI
2d ago

Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

The paper presents a new multi‑turn benchmark of 423 conversations with 1,661 labeled turn‑states to study when language models should refrain from answering. It shows that probes for unanswerability transfer well across datasets that share the same underlying signal, but fail to generalise to other forms of epistemic uncertainty. While a calibrated probe can identify underspecified turns more accurately than chance, it does not consistently improve overall generation quality compared to standard methods.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya
arXiv AI
2d ago

Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment

arXiv:2610.08482v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited,...

By Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, Xiaojuan Li, Mingrui Yang
arXiv AI
2d ago

Language-model ratings of depression reflect the rater more than the patient

The study examined how language‑model raters assess depression using the Patient Health Questionnaire across 880 raters and 189 interviews. Model choice accounted for 30% of symptom‑score variance, while stable participant differences explained 10.5%. Even raters with similar overall accuracy (AUC ≥ 0.70) disagreed on screening decisions for 40% of participants, and only recalibration with labeled data improved agreement and accuracy modestly.

By Baihan Lin
arXiv AI
2d ago

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

The paper argues that probe scores lack intrinsic meaning and should be interpreted relative to two reference points: a floor (what simple inputs predict) and a ceiling (what the full input predicts). The difference, called headroom, indicates the range where a probe can reveal that a model computes beyond what the input already provides. Experiments on transformers and real models show that headroom can vanish when the target no longer depends on hidden variables or when the input no longer reveals them, and that some previously claimed representations are largely explained by the input text alone.

By Pranjal Garg