arXiv Machine Learning By John S. H. Baxter, Pierre Jannin

Hallucinations and Constraints : Regulating surgical workflow recognition beyond accuracy

Read the original on arXiv Machine Learning →

arXiv:2608. 09332v1 Announce Type: new Abstract: Hallucinations are a major concern for the integration of artificial intelligence into medicine, although less explored in the realm of medical image processing.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 11

A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling

arXiv:2603. 27341v4 Announce Type: replace Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites.

By Kirill Skobelev, Eric Fithian, Yegor Baranovski, Jack Cook, Sandeep Angara, Shauna Otto, Zhuang-Fang Yi, John Zhu, Neeraj Mainkar, Margaux Masson-Forsythe, Daniel A. Donoho, X. Y. Han
Hugging Face Trending Papers
Jun 23

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models

Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented evaluation.

arXiv AI
Jun 2

Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training

arXiv:2601. 23220v2 Announce Type: replace-cross Abstract: Despite recent Multimodal Large Language Models (MLLMs)' linguistic prowess in medical diagnosis, we find even state-of-the-art MLLMs suffer from a critical perceptual deficit: geometric blindness.

By Anglin Liu, Ruichao Chen, Yi Lu, Hongxia Xu, Jintai Chen