arXiv Machine Learning By Anton Rasmussen, Hong Qin

Quantization Effects on Biomedical LLM Reliability

Read the original on arXiv Machine Learning →

arXiv:2608. 03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

The study evaluates how quantization affects accuracy and safety of five 7‑8B language models on clinical benchmarks. INT8 GPTQ shows minimal degradation (≤1.9%) across tasks, while INT4 causes substantial, model‑dependent drops, especially in high‑risk scenarios and safety metrics. Recovery methods such as clinical calibration substitution and QLoRA fine‑tuning yield mixed results, underscoring the need for task‑specific validation.

By Leonard Twagirayezu, Prasenjit Mitra
arXiv Machine Learning
Sep 11

Domain-Specific Hallucination Detection in Large Language Models

The paper introduces a multi‑signal pipeline for detecting hallucinations in large language models, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves high performance (F1 = 0.915, AUROC = 0.977) across QA, summarization, and dialogue, and shows that 25 % of training data yields 77 % of full‑data performance. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator cuts hallucination rates from 85.5 % to 37.7 %, and that domain‑specific fine‑tuning (PubMedBERT on SciFact) outperforms general‑domain models for biomedical text.

By Varun Teja Chundru, Debasmita Biswas