arXiv:2606. 11208v1 Announce Type: cross Abstract: Biomedical findings often seem to conflict across studies, but many of these differences are context-dependent rather than true contradictions.
By Elias Hossain, Sanjeda Sara Jennifer, Sabera Akter Bushra, Niloofar Yousefi
The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.
By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
The paper introduces Clinical Intent Extraction (CIE), a task that transforms fragmented clinical action annotations into complete structured records called Clinical Intent Representation (CIR). CIR decomposes each action into verb, type, coded target, timing, condition, request‑intent (aligned to HL7 FHIR) and modality, adding dimensions absent in prior datasets. By re‑expressing five heterogeneous corpora into CIR, the authors create CIRCA, a benchmark of 10,011 harmonized intents with human‑validated subsets, crosswalks, and a deterministic FHIR R4 mapper, and demonstrate that existing models perform poorly on the full task, highlighting the need for targeted development.
By Alexander Apartsin, Yehudit Aperstein
arXiv:2609.21859v1 Announce Type: new
Abstract: Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely o...
By Jiacheng Lin, Zifeng Wang, Zheng Chen, Erick Scott, Ziwei Yang, Fanyang Yu, Sheng Zhong, Jimeng Sun
arXiv:2606. 28353v1 Announce Type: cross Abstract: Linking FDA-approved medical devices to their underlying United States Patent and Trademark Office (USPTO) patents enables critical applications such as recall root-cause analysis, M&A-driven IP discovery, and technology trajectory mapping.
By Yang Qingqing, Liu Haijiang, Li Moyan
arXiv:2608. 14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute.
By Dipankar Sarkar
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
By Martin Murin
arXiv:2608.15382v2 Announce Type: replace
Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy r...
By Ummara Mumtaz, Aimen Noor, Awais Ahmed
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
By Koyar Afrasyab
arXiv:2607. 06802v1 Announce Type: cross Abstract: Open physiological corpora are heterogeneous: they use different sensors, labels, sampling rates, recording settings, and clinical endpoints.
By Dovy Paukstys
arXiv:2408. 13378v5 Announce Type: replace Abstract: Workflows in drug-target interaction (DTI) assessment require integrating heterogeneous data from predictive models, curated resources, and observations from experimental literature.
By Yoshitaka Inoue, Tianci Song, Xinling Wang, Rui Kuang, Tianfan Fu, Augustin Luna
GxP-Agent is a multi‑agent system that transforms clinical trial protocols into CDISC‑compliant datasets by encoding the regulatory workflow as a directed acyclic graph (DAG). Each node in the DAG represents a domain‑specific task executed by a worker agent with specialized skill context, validation gates, and conditional retry logic. On the CDISC‑Bench benchmark, GxP-Agent with Claude Sonnet 4.6 achieved a perfect 100 % structural match for 49 variables across 254 records, outperforming single‑agent and flat multi‑agent baselines and enabling weaker models like GPT‑4.1 to reach 59.2 % under the same DAG.
By Jaime Yan