When Calibration Depends on the Scoring Rule: Quantized Biomedical LLM Classification
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables.
The study compares two ways of obtaining predictions from language models fine‑tuned on customer behavior: scoring answer tokens directly versus generating a written rationale and then scoring the resulting answer. Across 13 model‑domain cells covering four retail tasks, scored readouts consistently rank outcomes more accurately than generated readouts, with an AUC improvement ranging from 1.5 to 14.5 points. The authors also find that a third readout—eliciting a probability before any verdict—improves calibration but only when outcome rates are represented in training, and they recommend using generated rationales for interpretability while relying on scored heads for ranking.
The study evaluates how quantization affects accuracy and safety of five 7‑8B language models on clinical benchmarks. INT8 GPTQ shows minimal degradation (≤1.9%) across tasks, while INT4 causes substantial, model‑dependent drops, especially in high‑risk scenarios and safety metrics. Recovery methods such as clinical calibration substitution and QLoRA fine‑tuning yield mixed results, underscoring the need for task‑specific validation.
arXiv:2609.21288v1 Announce Type: new Abstract: Surface electromyography (sEMG)-based silent speech interfaces are limited by cross-user variability and calibration burden. We study a limited-data se...
arXiv:2609.14825v1 Announce Type: cross Abstract: Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency to hallucinate. Retrieval augmentation...
arXiv:2509.19375v2 Announce Type: replace-cross Abstract: Large language models are increasingly used for clinical text classification, where overconfident misclassifications can directly affect pati...