arXiv AI By Felipe Ocampo Osorio, Sebasti\'an Andr\'es Cajas Ordo\~nez, Maximin Lange, Rafi Al Attrach, Sahil Kapadia, Zakaria Laouabdia Sellami, Angelo Antonio Talio, Leo Anthony Celi

Towards a Deterministic Math Solver for Clinical Language Models

Read the original on arXiv AI →

The paper explores a new approach to improve arithmetic reliability in clinical language models by having the model generate case‑specific Python code that a local executor runs deterministically, rather than performing arithmetic directly. Experiments on the MedCalc‑Bench Verified dataset show that this Program‑Solve interface yields modest gains for larger models (Qwen2.5‑32B) but not for smaller ones (Qwen2.5‑7B), and it does not replace the need for verified formulas or accurate variable extraction. The study highlights that adding an executor can help some open‑weight models but is not a universal solution.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 1

Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

arXiv:2606. 31630v1 Announce Type: new Abstract: Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test can still be \emph{statistically} wrong -- a Gaussian likelihood for heavy-tailed data, a Poisson for over-dispersed counts, an invalid prior support, or a pathological parameterization.

By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao