arXiv AI

RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences

arXiv Computation and Language
Sep 23

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI: A Narrative Review and Conceptual Framework

The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.

By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
arXiv Computation and Language
Sep 25

Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark

The paper introduces Clinical Intent Extraction (CIE), a task that transforms fragmented clinical action annotations into complete structured records called Clinical Intent Representation (CIR). CIR decomposes each action into verb, type, coded target, timing, condition, request‑intent (aligned to HL7 FHIR) and modality, adding dimensions absent in prior datasets. By re‑expressing five heterogeneous corpora into CIR, the authors create CIRCA, a benchmark of 10,011 harmonized intents with human‑validated subsets, crosswalks, and a deterministic FHIR R4 mapper, and demonstrate that existing models perform poorly on the full task, highlighting the need for targeted development.

By Alexander Apartsin, Yehudit Aperstein
arXiv AI
Aug 19

GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents

GxP-Agent is a multi‑agent system that transforms clinical trial protocols into CDISC‑compliant datasets by encoding the regulatory workflow as a directed acyclic graph (DAG). Each node in the DAG represents a domain‑specific task executed by a worker agent with specialized skill context, validation gates, and conditional retry logic. On the CDISC‑Bench benchmark, GxP-Agent with Claude Sonnet 4.6 achieved a perfect 100 % structural match for 49 variables across 254 records, outperforming single‑agent and flat multi‑agent baselines and enabling weaker models like GPT‑4.1 to reach 59.2 % under the same DAG.

By Jaime Yan