arXiv AI

Beyond One-shot: AI Agents for Learning in Field Experiments

arXiv:2606. 02458v1 Announce Type: new Abstract: Organizations routinely run experiments for A/B testing, yet the data generated from one experiment is underutilized to inform subsequent intervention design.

arXiv AI
Jun 30

An AI agent for treatment reasoning over a biomedical tool universe

arXiv:2606. 28692v1 Announce Type: new Abstract: Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy.

By Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik
arXiv AI
Sep 4

Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications

The paper reviews how Large Language Models (LLMs) are being adapted for medical reasoning, moving beyond single-step answers to systems that can systematically, transparently, and verifiably reason. It introduces a taxonomy of enhancement techniques, split into training-time methods such as supervised fine‑tuning and reinforcement learning, and test-time methods like prompt engineering and multi‑agent systems. The review examines their application across text, image, and code modalities in key clinical areas—diagnosis, education, and treatment planning—and tracks the shift in evaluation benchmarks from simple accuracy to more nuanced assessments of reasoning quality and visual interpretability.

By Zizhan Ma, Wenxuan Wang, Meidan Ding, Shiyi Zheng, Shengyuan Liu, Jie Liu, Jiaming Ji, Linlin Shen, Yixuan Yuan, Wenting Chen
arXiv AI
Sep 15

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

ClinAgent is a conversational system that uses a ReAct-based LLM agent to retrieve and synthesize clinical trial information from multiple sources such as ClinicalTrials.gov, PubMed, and a local dataset. The agent iteratively reasons over user queries, selects appropriate tools, and refines its actions to provide grounded, up-to-date responses in natural language across multi-turn interactions. Evaluation across three phases shows that DeepSeek (thinking mode) excels in planning quality while Gemini 3.0 Flash delivers the highest overall performance and expert ratings, demonstrating the promise of agentic AI for improving clinical trial data access.

By Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini
arXiv AI
Jun 16

MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition

arXiv:2512. 11682v2 Announce Type: replace Abstract: Therapeutic decision-making in clinical medicine constitutes a high-stakes domain in which AI guidance interacts with complex interactions among patient characteristics, disease processes, and pharmacological agents.

By Tim Cofala, Christian Kalfar, Jingge Xiao, Johanna Schrader, Michelle Tang, Wolfgang Nejdl
arXiv AI
Aug 10

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

arXiv:2608. 07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy.

By Valentin Li\'{e}vin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang
arXiv AI
Sep 7

Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent

The paper introduces MedTraj, a framework that constructs, evaluates, and optimizes multi‑step reasoning trajectories for medical AI agents. It parses each trajectory into observations, evidence, numbered steps, and a conclusion, scoring them on coherence, evidence support, hallucination, completeness, and traceability. Experiments on CareQA, PubMedQA, and CECMed show that incorporating quality‑weighted trajectory context improves reasoning coherence and correctness while significantly reducing hallucinations.

By Yunqi Zhu, Wensheng Zhang, Xuebing Yang