arXiv AI

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

arXiv:2606. 28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and substantial expert effort.

arXiv AI
Jun 30

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

arXiv:2606. 28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates the complete systematic review and meta-analysis (SR/MA) workflow -- from literature search through statistical analysis, manuscript generation, and quality assurance -- with mandatory human oversight at critical decision points.

By Hsieh-Ting Lin, Jiunn-Tyng Yeh
arXiv AI
Jun 24

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

arXiv:2605. 06177v2 Announce Type: replace Abstract: Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness and tool registry differ, and integrating a new model into a comparable evaluation surface costs weeks of model-specific engineering.

By Jinge Wu, Hongjian Zhou, Mingde Zeng, Jiayuan Zhu, Junde Wu, Jiazhen Pan, Ayush Noori, Sean Wu, Honghan Wu, Fenglin Liu, David A. Clifton
arXiv Computation and Language
Aug 27

Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation

The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.

By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
arXiv AI
Sep 1

From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction

arXiv:2608.28974v1 Announce Type: new Abstract: Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and r...

By Daniel Kang, Michelle Hu, Soorya Ram Shimgekar, Shayan Vassef, Yufan Wang, Anit Kumar Sahu, Munmun De Choudhury, Vedant Das Swain, Christian Poellabauer, Li Yan Khor, Koustuv Saha, Robert Wojciechowski, Elliot Kidd, Piyum Zonooz, Navin Kumar
arXiv AI
Aug 28

Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset

The study evaluates a standalone large language model (LLM) versus a four‑step agentic pipeline for generating explanations of ICU mortality predictions on the eICU Demo dataset. XGBoost achieved an AUROC of 0.855 and an AUPRC of 0.332. In a 38‑case explanation subset, the standalone LLM produced one explanation with outcome leakage, while the agentic pipeline produced none; among 14 overlapping SHAP cases, the standalone LLM had higher SHAP alignment and direction consistency, whereas the agentic pipeline showed better guideline grounding, value specificity, and plausibility.

By Di Zhu, Chen Xie, Haoyun Zhang, Zihan Wei, Ziwei Wang, Jiazhao Shi, Ziyu Wang, Qiyang Xie
arXiv AI
Jun 19

Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

arXiv:2606. 19345v1 Announce Type: cross Abstract: The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent.

By Zhyar Rzgar K. Rostam, M\'arta P\'entek, J\'anos Tibor Czere, Zsombor Zrubka, L\'aszl\'o Gul\'acsi, G\'abor Kert\'esz
arXiv AI
Sep 7

A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support

The paper introduces Debate-Mixture-of-Agents (DMoA), a multi‑agent framework that structures role‑based interactions to mimic iterative diagnostic reasoning in clinical settings. Evaluated on 297 rare disease cases and 1,719 challenging cases, DMoA outperformed a GPT‑4o baseline, improving most likely diagnosis accuracy by 10.21 percentage points and safety rate by 11.36 percentage points. Ablation studies and further analyses revealed that these gains stem from the structured workflow rather than merely adding more models or longer outputs, and that performance benefits are influenced by the chosen structure, base model strength, and token budget.

By Chang Xia, Leilei Ouyang, Huimin Wang, Yong Zhao, Kang Li
arXiv AI
Jun 11

Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark

arXiv:2602. 19502v2 Announce Type: replace Abstract: Agentic AI systems are increasingly capable of autonomous data science workflows, yet clinical prediction tasks demand domain expertise that purely automated approaches struggle to provide.

By Lalitha Pranathi Pulavarthy, Raajitha Muthyala, Aravind V Kuruvikkattil, Zhenan Yin, Rashmita Kudamala, Saptarshi Purkayastha