arXiv Machine Learning

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

The study evaluates how quantization affects accuracy and safety of five 7‑8B language models on clinical benchmarks. INT8 GPTQ shows minimal degradation (≤1.9%) across tasks, while INT4 causes substantial, model‑dependent drops, especially in high‑risk scenarios and safety metrics. Recovery methods such as clinical calibration substitution and QLoRA fine‑tuning yield mixed results, underscoring the need for task‑specific validation.

arXiv AI
Aug 26

Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models

Neurosymbolic Alignment couples a 7B clinical language model with a graph‑based physiological world model to score candidate responses using homeostatic constraints, multi‑hop plausibility, and drug‑interaction penalties. This training‑time framework drives iterative on‑policy updates and achieves a 90.8% Clinical Safety Score on the CSB benchmark, outperforming ORPO, GPT‑4 (5‑shot), and a self‑correction pipeline. Ablation studies show that the HGNN scoring and iterative training are the key contributors to the safety gains.

By Abdulhady Abas Abdullah, Erik Cambria, Milena Zivkovic