arXiv Machine Learning By Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta, Anik Pal Chowdhury

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

Read the original on arXiv Machine Learning →

arXiv:2607. 27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 21

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation

arXiv:2607. 16303v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on language priors or medical templates rather than truly attending to diagnosis-critical regions.

By Yunhang Qian, Jiaquan Yu, Jiawei Liu, Meng Wang, Hongwei Bran Li, Xiaobin Hu
arXiv AI
Sep 7

Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent

The paper introduces MedTraj, a framework that constructs, evaluates, and optimizes multi‑step reasoning trajectories for medical AI agents. It parses each trajectory into observations, evidence, numbered steps, and a conclusion, scoring them on coherence, evidence support, hallucination, completeness, and traceability. Experiments on CareQA, PubMedQA, and CECMed show that incorporating quality‑weighted trajectory context improves reasoning coherence and correctness while significantly reducing hallucinations.

By Yunqi Zhu, Wensheng Zhang, Xuebing Yang
arXiv AI
Sep 3

Untangling the Mechanisms of Misleading Context in Medical Question Answering

The paper investigates how misleading context—specifically fabricated evidence and bare assertions—affects large language models’ medical question‑answering performance. Experiments on MedMisBench show that models are more prone to adopt answers based on assertions than fabricated evidence, and that these misleading cues are often disclosed in reasoning traces but rarely in final responses. A monitor that reads open reasoning traces can detect most corrupted decisions, whereas monitoring only responses is less effective.

By Robin Linzmayer, No\'emie Elhadad