arXiv AI

MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition

arXiv:2512. 11682v2 Announce Type: replace Abstract: Therapeutic decision-making in clinical medicine constitutes a high-stakes domain in which AI guidance interacts with complex interactions among patient characteristics, disease processes, and pharmacological agents.

arXiv AI
Jun 30

An AI agent for treatment reasoning over a biomedical tool universe

arXiv:2606. 28692v1 Announce Type: new Abstract: Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy.

By Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik
arXiv Machine Learning
Aug 24

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

The paper introduces an LLM-as-a-Judge framework for evaluating the outputs of an agentic drug discovery assistant, ChatInvent, deployed at AstraZeneca. It defines four quality dimensions—Completeness, Relevancy, Structural Clarity, and Scope Adherence—alongside deterministic Tool Call Correctness checks, and validates the judge against five expert annotators. After optimizing the best-performing judge with few-shot demonstrations, alignment with human majority votes improves from 0.80 to 0.86, and the framework reveals that informal question phrasing does not degrade output quality.

By Emma Granqvist, Roc\'io Mercado, Samuel Genheden
arXiv AI
Sep 15

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

ClinAgent is a conversational system that uses a ReAct-based LLM agent to retrieve and synthesize clinical trial information from multiple sources such as ClinicalTrials.gov, PubMed, and a local dataset. The agent iteratively reasons over user queries, selects appropriate tools, and refines its actions to provide grounded, up-to-date responses in natural language across multi-turn interactions. Evaluation across three phases shows that DeepSeek (thinking mode) excels in planning quality while Gemini 3.0 Flash delivers the highest overall performance and expert ratings, demonstrating the promise of agentic AI for improving clinical trial data access.

By Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
Hugging Face Trending Papers
Jun 30

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment.

arXiv AI
Jul 1

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

arXiv:2606. 31179v1 Announce Type: new Abstract: As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications.

By Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, Maximilian Rokuss, Mingyu Lu, Timothy Ossowski, Juan Manuel Zambrano Chaves, Cliff Wong, Peniel Argaw, Yashna Hasija, Mu Wei, Wen-wai Yim, Qin Liu, Zilin Jing, Jason Entenmann, Naoto Usuyama, Tristan Naumann, Hoifung Poon
arXiv AI
Jun 24

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

arXiv:2605. 06177v2 Announce Type: replace Abstract: Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness and tool registry differ, and integrating a new model into a comparable evaluation surface costs weeks of model-specific engineering.

By Jinge Wu, Hongjian Zhou, Mingde Zeng, Jiayuan Zhu, Junde Wu, Jiazhen Pan, Ayush Noori, Sean Wu, Honghan Wu, Fenglin Liu, David A. Clifton